AI-generated summaries
Today's ML research,
without the noise.
Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.
54
Papers today
8h
Update frequency
7
Days of history
Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction
Theory
- Machine learning models can improve CHF prediction accuracy compared to traditional empirical methods.
- Tube-trained ML models show favorable transferability to rod bundle geometries.
- The local hybrid model outperformed traditional methods, indicating the effectiveness of hybrid approaches.
- This study provides one of the first large-scale assessments of ML-based CHF models in a production-level subchannel analysis environment.
Read more
Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction
Summary
This paper evaluates the performance of machine learning (ML)-based models for predicting critical heat flux (CHF) in nuclear thermal hydraulics, specifically within the CTF subchannel code applied to square rod bundles. Traditional methods for CHF prediction have relied on empirical correlations and lookup tables, primarily developed from tube data. The study utilizes the Electric Power Research Institute (EPRI) rod bundle CHF database to assess both pure and hybrid ML models, including local and semilocal formulations. The findings indicate that tube-trained ML models can effectively transfer to rod bundle applications, outperforming conventional methods in various geometries and operating conditions. Notably, the local hybrid lookup table model exhibited the best overall performance, while the semilocal pure ML model remained competitive. This research represents a significant step in validating ML-based CHF models for reactor-relevant geometries, highlighting their potential for enhancing reactor safety and performance analysis.
Methodology
The study involved evaluating ML-based CHF models deployed within the CTF subchannel code using the EPRI rod bundle CHF database. Both pure ML and hybrid residual correction models were analyzed in local and semilocal formulations. The performance of these models was compared against traditional CHF prediction methods, including the Bowring correlation and W-3 correlation, to assess their predictive accuracy across various geometries and operating conditions.
Results
The ML-based models, particularly the local hybrid LUT model, demonstrated superior performance in predicting CHF for rod bundles compared to traditional methods. The study found that even models trained solely on tube data could achieve substantial improvements in CHF prediction accuracy for rod bundles, indicating their robustness and applicability in reactor environments.
Implications
The findings suggest that ML-based approaches can significantly enhance the predictive capabilities of CHF models in nuclear thermal hydraulics, potentially leading to improved reactor safety and operational flexibility. This research paves the way for the integration of advanced ML techniques into production-level thermal hydraulic analysis tools.
Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning
Reinforcement Learning
Robotics
- Introduces a benchmark for evaluating muscle activation prediction accuracy in MIRL pipelines.
- Provides the first direct comparison between HyFyDy and MuscleMimic simulation environments.
- Evaluates the impact of simulator fidelity, subject-specific personalization, and model dimensionality on muscle activation predictions.
- Offers OpenSim-to-reference data and a trained subject-specific HyFyDy model for future benchmarking.
Read more
Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning
Summary
This paper presents a systematic comparison of two advanced motion-imitation reinforcement learning (MIRL) pipelines: HyFyDy/SCONE and MuJoCo/MyoSim. The authors aim to evaluate not only the kinematic accuracy of these pipelines but also their ability to replicate underlying neuromuscular behaviors, which is crucial for the design of robotic assistive devices. The study utilizes a common set of human motion-capture and electromyography (EMG) data to benchmark the muscle activation predictions of both frameworks. Results indicate that while both pipelines achieve similar kinematic performance, HyFyDy demonstrates superior alignment with experimental EMG data, suggesting its greater physiological realism. The paper emphasizes the need for further development in both frameworks to enhance their physiological accuracy and applicability in assistive robotics.
Methodology
The authors conducted three comparative studies involving both MIRL frameworks using equivalent musculoskeletal models. They assessed the accuracy of muscle activation predictions against synchronized motion-capture and EMG measurements, exploring variations in model complexity and personalization.
Results
The study found that HyFyDy produced muscle activation patterns that were more closely aligned with experimental EMG data, with RMSE and correlation values of (0.164, 0.4) compared to (0.344, 0.11) for MuJoCo. This indicates that HyFyDy is currently more suitable for musculoskeletal modeling, although both frameworks require further refinement to enhance physiological realism.
Implications
The findings suggest that HyFyDy's advanced physiological modeling could improve the design and control of robotic assistive devices. The benchmarks established in this study may facilitate the development of more effective controllers that generalize across different users and tasks.
Optimization Geometry of Equivalent Brownian RKHS Representations
Optimization
Theory
- Establishes a controlled finite Brownian RKHS with equivalent nodal, increment, and spectral coordinates.
- Demonstrates that nodal and spectral GD trajectories are identical under specific conditions.
- Derives a grid-resolution-independent bound for Brownian-regularized least squares.
- Introduces a maximality theorem for the standard Adam optimizer based on signed permutations.
Read more
Optimization Geometry of Equivalent Brownian RKHS Representations
Summary
This paper investigates the optimization geometry of equivalent finite parameterizations within a controlled finite Brownian Reproducing Kernel Hilbert Space (RKHS). The authors explore how different coordinate systems—namely nodal, increment, and spectral coordinates—can represent the same functions and norms while leading to distinct optimization algorithms. By employing classical finite-element methods, RKHS interpolation, and Brownian covariance identities, the study clarifies the relationships between these coordinate systems and their implications for optimization. The authors demonstrate that with identical initialization and step sizes, nodal and spectral gradient descent (GD) trajectories coincide, while increment GD behaves as an explicit Euler step under a constant Brownian/Sobolev metric. The paper also establishes bounds for Brownian-regularized least squares that are independent of grid resolution and introduces a maximality theorem for the standard Adam optimizer. Through numerical experiments, the authors validate their theoretical findings, confirming the equivalence of coordinate effects without altering the underlying functions or approximation spaces.
Methodology
The authors utilize a controlled finite Brownian RKHS framework to analyze the optimization geometry of different coordinate systems. They employ classical techniques from finite element analysis, RKHS interpolation, and Brownian covariance identities to derive relationships between the coordinate systems. The study includes theoretical derivations of optimizer trajectories and bounds, alongside numerical experiments to validate the findings.
Results
The main results include the identification of equivalent nodal, increment, and spectral coordinates within the Brownian RKHS framework, the demonstration of identical trajectories for nodal and spectral GD, and the derivation of a bound for Brownian-regularized least squares that is independent of grid resolution. Additionally, the authors present a maximality theorem for the standard Adam optimizer, confirmed through numerical tests.
Implications
The findings have significant implications for the design of optimization algorithms in machine learning, particularly in contexts where different parameterizations of models can lead to varying performance. Understanding the geometry of optimization in RKHS can enhance algorithm efficiency and robustness, potentially influencing future research in optimization techniques and their applications in various machine learning domains.
Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks
Reinforcement Learning
Optimization
- Introduces a shared latent-space framework for urban transportation calibration and control.
- Develops a combinatorial MLP-autoencoder for efficient simulator calibration.
- Implements a deep Q-learning agent for dynamic traffic optimization.
- Achieves up to 51% reduction in system-wide travel times in empirical tests.
Read more
Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks
Summary
This paper addresses the complex optimization challenges in urban transportation networks, focusing on the calibration of high-fidelity simulators and real-time operational control. The authors propose a shared latent-space framework that integrates simulator calibration with reinforcement learning control through a common learned representation of urban traffic dynamics. The framework utilizes a combinatorial MLP-autoencoder architecture to derive low-dimensional representations linking simulator inputs (like origin-destination demand and network parameters) to outputs (such as travel times and congestion patterns). This approach enhances sample efficiency compared to traditional methods, allowing for better fitting to observational data within fixed computational budgets. Additionally, a deep Q-learning agent is implemented to optimize dynamic traffic assignment via scheduling and routing adjustments. Empirical evaluations demonstrate that this method can reduce system-wide travel times by up to 51% compared to baseline operations. The shared latent representation not only aids in dimensionality reduction for Bayesian calibration but also enhances the reinforcement learning state representation, facilitating adaptive operational control. The findings underscore the potential of deep learning techniques in urban mobility planning, particularly for large-scale networks where conventional optimization methods encounter computational limitations.
Methodology
The methodology involves a two-part framework: first, a combinatorial MLP-autoencoder architecture is used to learn a low-dimensional latent representation of transportation simulator behavior for calibration. Second, a deep Q-learning agent employs this latent representation to optimize traffic control through scheduling and routing adjustments, enhancing both sample efficiency and computational tractability.
Results
The proposed framework significantly reduces system-wide travel times by up to 51% compared to baseline operations. The use of a shared latent representation improves the efficiency of both Bayesian calibration and reinforcement learning control, demonstrating superior sample efficiency and better fitting to observational data.
Implications
The findings suggest that deep learning methods can transform urban mobility planning and management, particularly in large-scale networks where traditional optimization techniques struggle. This integrated approach could lead to more adaptive and efficient transportation systems, enhancing urban mobility and reducing congestion.
Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain–Computer Interfaces
Multimodal
- Bio-MF enables low-latency EEG-to-fNIRS signal generation, achieving 857x speedup over traditional methods.
- The framework directly predicts clean fNIRS signals, improving fidelity and reducing non-physiological artifacts.
- Integration of advanced techniques like Spatial-Temporal Interactive 4D Encoding enhances performance across heterogeneous sensor layouts.
- Significant accuracy improvements in motor imagery classification are observed when using EEG + synthetic fNIRS data.
Read more
Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain–Computer Interfaces
Summary
This paper presents Bio-MF, a novel framework for generating functional near-infrared spectroscopy (fNIRS) signals from electroencephalography (EEG) data in hybrid motor-imagery brain-computer interfaces (MI-BCIs). The proposed method addresses the limitations of existing EEG-to-fNIRS generation techniques, which often suffer from high latency and require extensive pretraining. Bio-MF employs a latent-free one-step MeanFlow approach that directly predicts clean fNIRS signals, thus enhancing generation speed and fidelity. The framework integrates several innovative components, including Spatial-Temporal Interactive 4D Encoding and noise-level-gated FFT regularization, to maintain the integrity of hemodynamic structures across different sensor layouts. Experimental results demonstrate that Bio-MF significantly improves classification accuracy compared to EEG-only systems, achieving notable performance gains on two datasets. The framework generates fNIRS trials in just 7.0 ms, marking a substantial speedup over traditional methods. This advancement has the potential to enhance the usability and effectiveness of hybrid MI-BCIs in practical applications.
Methodology
Bio-MF utilizes a one-step MeanFlow framework that predicts fNIRS signals conditioned on EEG data. It incorporates Spatial-Temporal Interactive 4D Encoding to capture complex relationships and employs noise-level-gated FFT regularization to enhance signal quality. The model operates without iterative refinement, allowing for rapid generation of fNIRS signals.
Results
On two datasets, Bio-MF achieved accuracy improvements of 3.37 and 4.15 percentage points for HbR and HbO, respectively, compared to EEG-only systems. The framework generated fNIRS trials in 7.0 ms, demonstrating a significant reduction in latency compared to traditional methods.
Implications
The development of Bio-MF has significant implications for the design and implementation of hybrid MI-BCIs, making them more accessible and efficient for real-time applications in rehabilitation, assistive technologies, and human-machine interaction.
Toward individual-level calibration in affect recognition with perceptual adjustment queries
Computer Vision
- Introduces a framework for individual-level calibration in affect recognition using PAQs.
- Demonstrates that perceptual sensitivity varies significantly among individuals.
- Validates the effectiveness of PAQ calibration through behavioral measures in a 2AFC task.
- Shows significant improvements in perceived task difficulty and response time consistency.
Read more
Toward individual-level calibration in affect recognition with perceptual adjustment queries
Summary
This paper addresses the issue of perceptual miscalibration in affect recognition tasks, where identical stimuli may impose different perceptual difficulties across individuals due to variations in perceptual sensitivity. The authors propose a novel framework utilizing Perceptual Adjustment Queries (PAQs) to estimate each participant's Just Noticeable Difference (JND) along the facial affect spectrum. By employing a Two-Alternative Forced-Choice (2AFC) task, the framework normalizes perceptual difficulty, allowing for individualized calibration of affect recognition tasks. The validation of this approach demonstrates that PAQ calibration significantly equalizes perceived task difficulty among participants, reduces mean response time, and minimizes between-subject variance in response time. This establishes PAQ as a practical tool for enhancing individualized perceptual calibration in facial affect recognition, with implications for various applications in affective computing and human-machine interaction.
Methodology
The methodology involves a calibration paradigm based on Perceptual Adjustment Queries (PAQs), where participants identify a facial expression that is just-noticeably-different from a reference face. This is followed by a 2AFC task to assess their ability to distinguish between facial expressions. The study uses behavioral measures, including binary metacognitive difficulty judgments and response time variance decomposition, to validate the calibration framework.
Results
The results indicate that PAQ calibration significantly equalizes perceived task difficulty at an individual level compared to non-calibrated baselines and population-level Weibull calibration. Additionally, the mean response time decreased, and the variance in response time between subjects was reduced, highlighting the effectiveness of the PAQ approach.
Implications
The findings suggest that individualized perceptual calibration can enhance the accuracy of affect recognition systems, which is crucial for applications in mental health diagnostics, human-computer interaction, and personalized user experiences in various domains.
Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding
Audio & Speech
Multimodal
- Multi-subject pretraining significantly outperforms zero-shot transfer and direct fine-tuning methods.
- Only 3 minutes of target-subject calibration can yield comparable performance to longer calibration sessions.
- Increasing the number of pretraining subjects leads to substantial improvements in decoding accuracy.
- A subject-specific MLP adapter did not provide any detectable benefit in personalization.
Read more
Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding
Summary
This paper addresses the challenges of cross-user variability and calibration burden in surface electromyography (sEMG)-based silent speech interfaces. The authors investigate a limited-data scenario where 27 participants contributed less than 0.5 hours of sEMG data each. They propose a methodology that involves multi-subject pretraining followed by short calibration for personalized speech decoding within a closed 50-sentence corpus. The study employs a leave-one-subject-out evaluation approach, initializing from a single-subject checkpoint and fine-tuning on the target participant's data. The results demonstrate that this pipeline significantly reduces character and word error rates compared to traditional methods, achieving a character error rate (CER) of 21.7% and a word error rate (WER) of 31.9%. The findings highlight the effectiveness of leveraging multi-subject pretraining to enhance personalization with minimal calibration data, suggesting a promising direction for developing efficient silent speech interfaces.
Methodology
The authors utilized a leave-one-subject-out evaluation strategy, where they pretrained models on data from multiple subjects and fine-tuned them on a target subject's data. They analyzed the impact of the number of pretraining subjects and the amount of calibration data on performance metrics such as character error rate (CER) and word error rate (WER). The study involved a closed 50-sentence corpus and used a convolutional front-end with a transformer encoder for decoding.
Results
The proposed method achieved a CER of 21.7% and a WER of 31.9% with short calibration, compared to 49.3% CER without calibration and 68.0% CER for direct fine-tuning. The macro-averaged CER improved from 74.4% with one pretraining participant to 21.7% with 26 participants. Calibration with just 3 minutes of target-subject data resulted in a CER of 20.5%, showing no significant difference from longer calibration sessions.
Implications
The findings suggest that short-calibration personalization can be effectively implemented in sEMG-based silent speech interfaces, making them more accessible and user-friendly. This approach could facilitate the development of assistive communication technologies for individuals with speech impairments, reducing the need for extensive training data.
Bilevel Optimization of Topology and Hyperparameters (BOTH)
Optimization
- Introduces a bilevel optimization framework for simultaneous tuning of hyperparameters in topology optimization.
- Utilizes automatic differentiation to compute hypergradients, enabling efficient hyperparameter optimization.
- Demonstrates scalability to thousands of hyperparameters with computational costs comparable to standard TO runs.
- Shows effectiveness on stress-constrained and compliance problems, improving upon traditional heuristic methods.
Read more
Bilevel Optimization of Topology and Hyperparameters (BOTH)
Summary
This paper addresses the challenges of hyperparameter tuning in topology optimization (TO), which is crucial for automating design processes. The authors propose a novel approach that utilizes automatic differentiation to derive 'hypergradients' for simultaneous tuning of hyperparameters alongside the primary optimization. This method allows for the evaluation of only a few steps of TO, significantly reducing the computational burden while maintaining effectiveness. The authors demonstrate the efficacy of their approach on stress-constrained and compliance problems, showcasing its ability to handle thousands of hyperparameters efficiently. The results indicate that this bilevel optimization framework can outperform traditional methods that rely on heuristic tuning, thus providing a more systematic and principled way to optimize hyperparameters in TO workflows.
Methodology
The authors employ a bilevel optimization strategy where the upper-level optimization iteratively proposes hyperparameter configurations using surrogate-based methods, while the lower-level optimization solves the TO problem from scratch for each proposed configuration. Automatic differentiation is used to compute hypergradients, facilitating the tuning of hyperparameters in tandem with the primary optimization.
Results
The proposed method successfully reduced the rate of failed or poorly converged designs in topology optimization tasks. It demonstrated that evaluating just one or two iterations of TO provides sufficient information for effective hyperparameter tuning, achieving results comparable to traditional methods with significantly fewer computational resources.
Implications
This work has the potential to streamline the design process in engineering and materials science by automating hyperparameter tuning in topology optimization. It can lead to more efficient design workflows, reducing reliance on expert intuition and trial-and-error methods, and enabling broader adoption of TO in various applications.
Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals
Large Language Models
Efficient ML
Optimization
- Introduces a decomposition of quantization error into AGWC and orthogonal residual components.
- Clarifies the roles of weight compensation and transformation design in addressing quantization errors.
- Provides theoretical guidance for randomized rotation, sign selection, and channel scaling.
- Demonstrates competitive performance of backpropagation-free quantization methods against gradient-trained approaches.
Read more
Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals
Summary
This paper addresses the challenges of post-training weight-activation quantization (PTQ) in large language models (LLMs), particularly the difficulties associated with aggressive 4-bit quantization (W4A4) due to activation outliers. The authors propose a novel framework that decomposes the local quantization error into two components: an activation-guided weight compensation (AGWC) term and an orthogonal residual. This decomposition clarifies which error components can be mitigated through weight optimization and which require transformation design. The authors analyze the orthogonal residual to provide theoretical guidance for techniques such as randomized Hadamard rotation, sign selection, and channel scaling. They derive practical guidelines for these methods, demonstrating how random signs can suppress interference from outlier channels and how second-moment balancing leads to effective scaling rules. The proposed methods are evaluated across eight Llama and Mistral models, achieving performance comparable to gradient-trained methods without requiring backpropagation. Overall, the paper contributes a deeper understanding of quantization error components and offers practical strategies for improving quantization in LLMs.
Methodology
The authors develop a local error-decomposition framework to analyze weight-activation quantization error. They derive bounds for the orthogonal residual based on activation outliers and regular activation quantities. The framework is used to guide the application of randomized Hadamard rotation, sign selection, and channel scaling in practical quantization configurations.
Results
The proposed methods were evaluated on eight Llama and Mistral models, achieving performance that is competitive with gradient-trained SpinQuant configurations, demonstrating the effectiveness of the theoretical guidelines in practical applications.
Implications
The findings provide a structured approach to improving quantization techniques in large language models, potentially leading to more efficient deployment of LLMs in resource-constrained environments. The insights into error components can inform future research and development in quantization strategies.
Routine Blood Tests Outperform CRP for Distinguishing Bacterial From Viral Infection in Children
Theory
- Routine blood tests, including CBC, provide better predictive capabilities than CRP alone for distinguishing between bacterial and viral infections in children.
- The XGBoost model outperformed the CRP-only decision rule, achieving an AUC of 81.7%.
- Incorporating multiple blood parameters can enhance diagnostic accuracy and inform antibiotic prescribing practices.
- The study emphasizes the need for pediatricians to consider a broader range of laboratory results when diagnosing infections.
Read more
Routine Blood Tests Outperform CRP for Distinguishing Bacterial From Viral Infection in Children
Summary
This study addresses the challenge of distinguishing between bacterial and viral infections in children, a critical task for appropriate antibiotic prescribing. The authors conducted a retrospective analysis of 906 pediatric patients aged 2-14 years who were evaluated for suspected infections at a pediatric hospital in Bulgaria. The study compared the predictive capabilities of routine blood tests, specifically Complete Blood Count (CBC) and C-reactive protein (CRP), against the traditional CRP-only approach. The findings revealed that machine learning models, particularly XGBoost, significantly outperformed the CRP baseline model in distinguishing infection types. The best-performing model achieved an AUC of 81.7%, indicating a robust ability to differentiate between bacterial and viral infections. The study highlights the importance of incorporating multiple blood parameters, such as white blood count and lymphocyte count, into clinical decision-making to reduce unnecessary antibiotic prescriptions and combat antimicrobial resistance.
Methodology
The study utilized a retrospective analysis of data from 906 pediatric patients with confirmed viral or bacterial infections. Various supervised classification models, including logistic regression and XGBoost, were trained using features from CBC and CRP tests. Model performance was evaluated based on AUC, sensitivity, and specificity.
Results
The XGBoost model, which included all relevant features, achieved an AUC of 81.7%, with sensitivity at 70.8% and specificity at 79.2%. The model without CRP showed slightly lower performance but still outperformed the CRP-only baseline model.
Implications
The findings suggest that pediatricians should rely on a combination of blood test results rather than CRP alone to make informed decisions about antibiotic prescriptions. This approach could help mitigate the issue of antibiotic overprescription and combat antimicrobial resistance.
GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning
Reinforcement Learning
Robotics
Optimization
- GEM-MPC combines exploitation and exploration in planning through a dual-policy framework.
- Introduces Gated Prior Distillation to filter stale planning data efficiently.
- Demonstrates consistent performance improvements over existing methods in continuous control tasks.
- Reduces computational costs associated with reanalyzing planning distributions.
Read more
GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning
Summary
The paper addresses the challenge of effective exploration in high-dimensional continuous control within reinforcement learning (RL). It introduces GEM-MPC, a method that combines Model Predictive Path Integral (MPPI) control with learned policies to enhance the interaction between planning and learning. The authors identify issues with traditional planning methods, such as the misalignment between learned sampling policies and planner behavior, which can degrade performance. GEM-MPC employs a dual-policy approach: one policy clones the planner for exploitation, while the other is KL-regularized to promote exploration around the planner's actions. Additionally, the paper presents Gated Prior Distillation, a mechanism that selectively learns from stored planning distributions only when they offer better targets than current priors, thus mitigating the impact of stale data without incurring the high computational costs of full reanalysis. The proposed method demonstrates consistent improvements over existing planning-based RL baselines across various continuous-control benchmarks, achieving better performance with lower computational budgets.
Methodology
GEM-MPC utilizes MPPI to integrate two distinct policies: a greedy policy that clones the planner's behavior and an exploratory policy that is trained using KL-regularized reinforcement learning. The Gated Prior Distillation mechanism selectively distills planning distributions to ensure that only beneficial targets are used for training, thus avoiding the pitfalls of stale data.
Results
The results show that GEM-MPC consistently outperforms existing planning-based reinforcement learning methods under lower computational budgets across various continuous-control benchmarks, indicating its effectiveness in balancing exploration and exploitation.
Implications
GEM-MPC has potential applications in robotics and other domains requiring efficient exploration in high-dimensional action spaces, improving the performance of RL agents in complex environments.
Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability
Generative Models
Theory
Interpretability
- Introduction of Riemannian Neural Hamiltonian Flows (RNHF) for generative modeling on curved spaces.
- RNHF combines Riemannian kinetic energy with a learned scalar potential and uses a geodesic integrator.
- The model is symplectic, reversible, and volume-preserving, addressing limitations of traditional Hamiltonian flows in non-Euclidean settings.
- The framework provides insights into the interpretability of the learned Hamiltonian, linking potential forces to momentum distributions.
Read more
Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability
Summary
This paper introduces Riemannian Neural Hamiltonian Flows (RNHF), a novel generative model that extends the concept of Hamiltonian normalizing flows to Riemannian manifolds. Traditional Hamiltonian flows are effective in Euclidean spaces but struggle with the complexities of curved spaces. RNHF combines a fixed kinetic energy derived from the Riemannian metric with a learned scalar potential and employs a geodesic leapfrog integrator to ensure that the resulting transformations are symplectic, reversible, and volume-preserving. The paper also explores the interpretability of the learned Hamiltonian, demonstrating that the potential force competes with the momentum distribution pressure, leading to an implicit profile that can be matched to the target distribution. The authors provide a theoretical framework that elucidates how the learned potential can be interpreted, particularly in the context of isotropic Gaussian distributions. Numerical experiments validate the effectiveness of RNHF across various geometries, showing competitive sample quality and computational efficiency compared to existing Riemannian continuous normalizing flows.
Methodology
The authors developed RNHF by integrating the fixed kinetic energy from Riemannian geometry with a learned scalar potential. They utilized a geodesic leapfrog integrator to maintain symplectic properties and volume preservation. The model was trained using an evidence lower bound, avoiding the need for Jacobian determinants or divergence estimates, which are typically required in flow-based models.
Results
The RNHF model showed competitive sample quality and numerical efficiency in generating samples from distributions on Euclidean, hyperbolic, and spherical spaces. The interpretability of the learned potential was confirmed through theoretical analysis and numerical validation, demonstrating how the implicit profile aligns with the target distribution.
Implications
RNHF has the potential to enhance generative modeling in various applications where data is naturally represented in non-Euclidean spaces, such as spherical or hyperbolic data distributions. Its interpretability framework could also aid in understanding complex generative processes in machine learning.
Generative inversion for early ranking of competing geologic interpretations
Generative Models
Multimodal
Theory
- The framework translates geological interpretations into testable geologic maps.
- It ranks competing interpretations based on their consistency with limited hydraulic-head observations.
- The method was validated through synthetic benchmarks and real-field case studies.
- It provides a quantitative compatibility score for each interpretation, aiding in decision-making.
Read more
Generative inversion for early ranking of competing geologic interpretations
Summary
This paper addresses the challenge of evaluating competing geological interpretations in subsurface modeling, particularly when data is scarce. The authors propose a novel framework that converts textual geological descriptions into geologic maps, which can be tested against existing hydraulic-head observations. By utilizing a text-to-image foundation model, the framework generates an ensemble of geological images for each interpretation. A variational autoencoder creates a latent representation specific to each interpretation, while a supervised inverse network maps hydraulic-head observations into this latent space. The framework then simulates steady-state flow to predict hydraulic heads, comparing these predictions to actual measurements to derive a compatibility score for each interpretation. The method was tested on a synthetic benchmark and real-world case studies, demonstrating its effectiveness in ranking interpretations based on their consistency with observed data. The results indicate that the framework can reliably identify the most accurate geological interpretation, potentially saving significant costs in subsurface decision-making.
Methodology
The authors developed a workflow that includes generating geological images from text descriptions using a foundation model, creating interpretation-specific latent representations with a variational autoencoder, and mapping hydraulic-head observations into this latent space using a supervised inverse network. The framework then simulates flow to predict hydraulic heads and calculates compatibility scores based on the mismatch between predicted and observed values.
Results
In tests involving 925 cases, the framework showed a mean RMSE increase from 0.197 for the most accurate interpretation to 0.280 for the least accurate. In a real-world application to the Culebra Dolomite Member, the revised model received a compatibility weight of 0.991, significantly higher than the original model's weight of 0.009, aligning with independent evidence.
Implications
This framework has the potential to transform subsurface decision-making by providing a systematic, data-driven approach to evaluate geological interpretations early in the investigation process. It can significantly reduce costs associated with drilling and improve the accuracy of subsurface models in various applications, including groundwater management, hydrocarbon production, and geologic carbon storage.
From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities
Multimodal
Time Series
- Comparative evaluation of three deep learning architectures for emotion recognition.
- Multimodal sensing configurations consistently outperform single modality setups.
- Transformer model achieved the highest accuracy on WESAD dataset.
- LSTM model performed best on EmoWear for arousal and valence recognition.
Read more
From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities
Summary
This paper presents a comprehensive study on physiological emotion recognition using multimodal wearable sensors, addressing the limitations of previous research that often focused on single models or datasets. The authors evaluate three deep learning architectures—Bidirectional LSTM, Temporal Convolutional Network (TCN), and Transformer—across two datasets (WESAD and EmoWear) under various sensing configurations (wrist-only, chest-only, and multimodal). A unified preprocessing pipeline and participant-independent leave-one-subject-out cross-validation (LOSO-CV) are employed to ensure consistency. The study also explores soft-voting ensembles, classical machine learning baselines, and conducts a systematic sensor ablation study. The findings reveal that the Transformer model achieves the highest accuracy on the WESAD dataset, while the LSTM excels on EmoWear for arousal and valence recognition. The results indicate that the performance of architectures is dataset-dependent, and multimodal configurations consistently outperform single modality setups. Additionally, a sampling frequency of 4 Hz is identified as a practical choice, balancing performance and computational efficiency. This work provides valuable insights for selecting architectures, sensing modalities, and sampling frequencies in wearable physiological emotion recognition.
Methodology
The study employs a comparative analysis of Bidirectional LSTM, TCN, and Transformer models using two multimodal datasets (WESAD and EmoWear). A unified preprocessing pipeline and LOSO-CV are utilized. The authors also implement soft-voting ensemble strategies and conduct systematic sensor ablation studies to assess the impact of different configurations and sampling frequencies on recognition performance.
Results
The Transformer achieved the highest multimodal accuracy on the WESAD dataset (99.02% ± 0.51%), while the LSTM outperformed on EmoWear for arousal (91.80% ± 1.06%) and valence (89.96% ± 0.36%). Multimodal configurations consistently provided better performance than wrist-only or chest-only setups across all architectures. The analysis indicated that a sampling frequency of 4 Hz is effective, offering comparable performance to higher frequencies with reduced training costs.
Implications
The findings suggest that selecting appropriate deep learning architectures and sensor modalities can significantly enhance the effectiveness of physiological emotion recognition systems. This research has implications for mental health monitoring, affective computing, and the development of adaptive human-computer interfaces.
Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
NLP
Large Language Models
Efficient ML
- Introduces Elastic Threshold Attention (ETA) for efficient long-context decoding.
- Utilizes dynamic, contextual thresholds to optimize attention allocation.
- Employs multiplicative suppression of logits to maintain model quality during training.
- Achieves significant speed improvements in inference with a custom Triton decode kernel.
Read more
Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
Summary
The paper presents Elastic Threshold Attention (ETA), a novel architecture designed to enhance long-context decoding in Transformer models by addressing the memory-bandwidth bottlenecks associated with massive Key-Value (KV) caches. Traditional sparse attention methods often compromise model quality due to rigid heuristics that drop essential context. ETA overcomes this limitation by predicting dynamic, contextual thresholds from query representations, enabling the model to allocate dense-like context for complex retrieval tasks while pruning unnecessary tokens. The training methodology involves multiplicatively suppressing sub-threshold logits instead of outright deletion, which maintains a smooth attention distribution and prevents representation collapse. This approach allows ETA to rival dense attention models in various tasks, achieving approximately 85% training sparsity and 38% active decode density. The implementation of a custom decode kernel in Triton facilitates rapid KV block screening, yielding up to 2.5× speed improvements over existing methods like FlashAttention-2. Additionally, an offline calibration algorithm is introduced for domain-specific applications, further optimizing attention computation by 27%. Overall, ETA demonstrates a significant advancement in balancing speed and quality in long-context decoding tasks.
Methodology
ETA employs a lightweight linear predictor to generate dynamic thresholds based on query representations. During training, it suppresses sub-threshold logits multiplicatively, creating a uniform attention floor that prevents the formation of attention sinks. At inference, a custom Triton kernel screens KV blocks efficiently, allowing for rapid decoding without excessive memory transfers.
Results
The ETA model, with 1.45 billion parameters, matches the performance of dense attention models across various tasks, including language modeling and commonsense reasoning, while achieving about 85% training sparsity and 38% active decode density. The custom decode kernel provides up to 2.5× speedup in decoding time compared to FlashAttention-2.
Implications
ETA's approach to contextual sparsity and efficient decoding can significantly enhance the performance of large language models in real-time applications, particularly in scenarios requiring long-context processing. Its methods may be applicable in various domains, including natural language processing and machine learning inference optimization.
Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers
Efficient ML
Optimization
Theory
- Introduces a decoupled, layerwise method for structured sparsification of neural networks.
- Proves the equivalence of the decoupled objective to a joint penalty at optimality.
- Demonstrates a wider usable range of regularization strength with lower rates of catastrophic over-pruning.
- Validates the method through controlled experiments and ablation studies.
Read more
Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers
Summary
This paper introduces a novel method for structurally sparsifying fully connected layers in pretrained neural networks through a decoupled, layerwise approach. Unlike traditional methods that penalize all layers simultaneously, the proposed method focuses on shallow two-layer subnetworks, normalizing inner weights and applying a structured group penalty to outer weights sequentially. This decoupling leads to a more robust optimization process, allowing for a wider range of regularization strengths and reducing the risk of catastrophic over-pruning while maintaining comparable accuracy to joint penalization methods. The authors establish the theoretical equivalence of their decoupled formulation to a specific joint penalty at optimality for any positively homogeneous activation function. Through extensive numerical experiments, including classification tasks and stress tests on high-dimensional Physics-Informed Neural Networks (PINNs) and the OPT-1.3B model, the authors demonstrate the effectiveness of their approach in achieving efficient model compression without significant performance loss.
Methodology
The authors propose a decoupled layerwise method that processes shallow two-layer subnetworks sequentially. Inner weights are normalized, and a structured group penalty is applied to outer weights. The method incorporates fine-tuning into the sparsification loop, allowing for efficient pruning of neurons in fully connected layers of pretrained models. Theoretical proofs establish the equivalence of the decoupled and joint penalization approaches.
Results
The proposed method exhibits a wider range of effective regularization strengths and a significantly lower rate of catastrophic over-pruning compared to traditional joint penalization methods. The experiments confirm that the decoupled approach maintains comparable accuracy while enabling faster convergence and better performance in various tasks, including classification and sparse recovery.
Implications
This research has significant implications for model compression in resource-constrained environments, particularly for large-scale neural networks. The decoupled approach can facilitate the deployment of efficient models without compromising performance, making it suitable for applications in mobile and edge computing.
Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods
Optimization
Theory
Efficient ML
- Develops a theoretical framework for matrix-aware adaptive optimization using Online Mirror Descent.
- Introduces Row-wise and Column-wise Matrix AdaGrad algorithms that exploit matrix structure.
- Establishes tighter regret bounds compared to traditional AdaGrad for structured gradients.
- Demonstrates improved optimization stability and performance in experiments with matrix factorization and deep learning.
Read more
Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods
Summary
This paper addresses the limitations of existing adaptive optimization methods like AdaGrad and Adam, which are primarily designed for vector-valued parameters and do not leverage the matrix structure inherent in many machine learning models. The authors propose a novel Online Mirror Descent (OMD) framework that incorporates adaptive proximal functions specifically for matrix-valued parameters. By introducing row-wise and column-wise matrix proximal functions, they derive two new algorithms: Row-wise Matrix AdaGrad (Row-AdaGrad) and Column-wise Matrix AdaGrad (Column-AdaGrad). These algorithms adaptively scale learning rates based on accumulated row-wise or column-wise gradient norms, respectively. The theoretical foundation is established through online regret minimization, leading to tighter regret guarantees compared to traditional entry-wise AdaGrad. Experimental results demonstrate that these matrix-aware methods improve optimization stability and allow for larger learning rates and deeper networks, showcasing their effectiveness in matrix factorization and deep neural network training.
Methodology
The authors develop a general Online Mirror Descent framework with adaptive proximal functions tailored for matrix-valued parameters. They define row-wise and column-wise proximal functions based on squared matrix Mahalanobis norms, allowing for a principled derivation of adaptive optimization methods. The framework is analyzed through online regret minimization, leading to the formulation of Row-AdaGrad and Column-AdaGrad algorithms.
Results
The proposed algorithms, Row-AdaGrad and Column-AdaGrad, exhibit tighter regret guarantees than entry-wise AdaGrad when applied to structured gradients. Experimental evaluations show that these methods enhance optimization stability and enable training with larger learning rates and deeper networks, outperforming traditional adaptive optimizers in matrix factorization and deep neural network tasks.
Implications
The findings suggest that leveraging matrix structure in adaptive optimization can significantly improve training efficiency and stability in deep learning models. This work opens avenues for further research into matrix-aware optimization techniques, potentially benefiting a wide range of applications in machine learning.
Do Quantum Models Scale Like LLMs?
Theory
Large Language Models
Generative Models
- RydbergGPT demonstrates scaling laws similar to those of large language models near critical points.
- The loss function exhibits a power-law behavior with a loss floor correction in critical regions.
- Statistical structures of quantum measurement data can be compared to natural language corpora.
- Scaling behavior is influenced by the statistical structure of the training data, suggesting a broader applicability beyond natural language.
Read more
Do Quantum Models Scale Like LLMs?
Summary
This paper investigates the neural scaling laws of RydbergGPT, an autoregressive transformer model trained on qubit projective measurement data from interacting Rydberg atom arrays. The authors explore whether quantum models exhibit similar scaling behavior to large language models (LLMs) as the size of the training dataset increases. They find that near a critical point in the quantum system, the transformer loss follows a power-law with a loss floor correction, indicating a predictable scaling behavior. However, this power-law description diminishes when moving away from criticality. The authors compare the statistical structure of Rydberg measurements with natural language corpora using a mutual information 'two-point' function and discover that near-critical statistics align closely with those of natural language, while configurations far from criticality show a rapid decay in two-point functions. This suggests that multi-scale dependence plays a role in stable neural scaling, proposing that scaling behavior is a characteristic of the model-data interaction rather than being exclusive to natural language.
Methodology
The authors utilized RydbergGPT, an autoregressive transformer model, to learn from synthetic measurement samples generated by quantum Monte Carlo simulations. They analyzed the loss function's behavior as a function of training dataset size and compared the statistical properties of the measurement data with natural language using a mutual information 'two-point' function.
Results
The study found that the transformer loss near the critical point follows a power-law with a loss floor correction, while the scaling behavior is less predictable away from criticality. The mutual information analysis revealed that the statistical structure of near-critical measurements is similar to that of natural language, indicating a potential commonality in scaling behavior across different domains.
Implications
The findings suggest that quantum models can be effectively scaled in a manner akin to LLMs, which could lead to more efficient training and application of quantum machine learning models. This research opens avenues for further exploration of scaling laws in various structured probability distributions beyond natural language.
IncentRL: The Trade-Off Between Preference Guidance and Task Performance
Reinforcement Learning
Theory
Optimization
- IncentRL introduces a KL penalty to balance preference guidance and task performance in RL.
- The framework establishes theoretical bounds for maintaining the original optimal policy despite preference shaping.
- Empirical results show a significant increase in task success rates with preference guidance compared to a baseline.
- The study highlights the importance of managing the trade-off between additional guidance and task distortion.
Read more
IncentRL: The Trade-Off Between Preference Guidance and Task Performance
Summary
The paper presents IncentRL, a framework designed to incorporate preference-based reward shaping in reinforcement learning (RL) while maintaining the integrity of the original task objective. The authors highlight the potential pitfalls of adding preference signals to rewards, which can inadvertently alter the task being optimized. IncentRL introduces a Kullback-Leibler (KL) penalty that measures the divergence between a specified outcome distribution and a preferred distribution. The authors derive theoretical bounds for external-value perturbation and establish conditions under which the original optimal policy can be preserved. They also explore the implications of large preference weights on the learning process. The practical implementation of IncentRL is evaluated using the MiniGrid DoorKey-8x8 environment, demonstrating a significant improvement in success rates compared to a baseline without preference shaping. The results suggest that while preference guidance can enhance learning, it must be carefully balanced to avoid distorting the task objective. The paper concludes with a discussion of the trade-offs involved in preference-based RL and outlines future research directions to further isolate the effects of KL shaping.
Methodology
The authors derive theoretical results for IncentRL within the context of finite discounted Markov decision processes (MDPs). They analyze the implications of adding a KL penalty to the reward structure, focusing on the conditions necessary for preserving optimal policies. A practical implementation is tested in the MiniGrid environment, utilizing a distance-based outcome proxy and a fixed preference distribution.
Results
The implementation of IncentRL in the MiniGrid DoorKey-8x8 environment achieved a mean success rate of 98% after two million training steps with a coefficient of 0.01, compared to 90.5% for the zero-coefficient baseline. The results indicate that preference guidance can significantly improve learning outcomes while maintaining the original task's integrity.
Implications
The findings suggest that preference-based reward shaping can be effectively integrated into RL frameworks, potentially leading to more efficient learning strategies in complex environments. The insights gained from this study could inform the design of RL systems that require nuanced guidance without compromising task objectives.
LLMs as Feature Engineers for Text-and-Tabular Prediction
NLP
Large Language Models
Interpretability
- Introduces an iterative framework for automating feature extraction from unstructured text.
- Demonstrates that LLM-generated features significantly enhance predictive performance when combined with traditional methods.
- Ensures instance-level interpretability, providing a clear semantic audit trail for predictions.
- Utilizes an error-driven feedback mechanism to optimize feature generation efficiently.
Read more
LLMs as Feature Engineers for Text-and-Tabular Prediction
Summary
This paper presents an innovative framework that automates the extraction of interpretable categorical features from unstructured text for use in tabular prediction models. The proposed method utilizes two large language models (LLMs): a generator LLM that proposes semantic definitions for features and an extractor LLM that applies these definitions to annotate datasets. A downstream tabular model evaluates the predictive performance of the generated features. The framework incorporates an error-driven feedback mechanism that translates explicit model errors into natural language, guiding the generator LLM to refine its feature proposals. This iterative process accelerates feature discovery by up to three times compared to unguided searches. The authors demonstrate that the generated features exhibit strong complementarity with traditional representations like TF-IDF and dense embeddings, leading to improved predictive performance. Additionally, the framework ensures instance-level interpretability, as the discovered features rank highly in SHAP importance and provide a transparent audit trail for predictions.
Methodology
The methodology involves an iterative loop where a generator LLM proposes feature definitions, an extractor LLM annotates the dataset based on these definitions, and a downstream tabular model evaluates the features. The framework is optimized using natural language feedback derived from explicit model errors, guiding the LLM to improve feature generation iteratively.
Results
The framework was evaluated on three public datasets, showing that the error-driven approach accelerates feature discovery by up to 3× compared to unguided searches. The generated features were found to be complementary to existing representations, outperforming any subset when combined with TF-IDF and dense embeddings. The LLM-generated features also dominated SHAP importance rankings, confirming their interpretability.
Implications
This framework has significant implications for tasks that require the integration of unstructured text with tabular data, such as recommendation systems and click-through rate prediction. It allows for efficient feature engineering while maintaining interpretability, which is crucial for model transparency and trust.
Trading Depth for Time in Recurrent Transformers
NLP
Large Language Models
Efficient ML
- Introduces Latent Recurrent Transformers (LRTs) that utilize latent thought tokens for hidden-state refinement.
- Demonstrates that temporal computation can recover significant performance improvements with fewer parameters compared to increased physical depth.
- Shows that one thought token can achieve 67% and 81% of the performance gains from doubling the model depth in different architectures.
- Highlights the importance of feedback connections in enhancing the performance of recurrent Transformers.
Read more
Trading Depth for Time in Recurrent Transformers
Summary
This paper investigates the trade-off between increasing computational depth and adding temporal steps in Recurrent Transformers, specifically through the introduction of Latent Recurrent Transformers (LRTs). LRTs utilize a latent thought token that refines the hidden state between vocabulary tokens, allowing for a comparison of the benefits of temporal computation versus physical depth. The authors demonstrate that by inserting a thought token, they can achieve significant performance improvements while using fewer parameters compared to a model with increased physical depth. The study reveals that one thought token can recover 67% and 81% of the performance gains from doubling the depth in two different model architectures, while using approximately 48% fewer parameters. This suggests that temporal computation can be a more parameter-efficient strategy in recurrent architectures, offering a new perspective on optimizing Transformer models.
Methodology
The authors compare two approaches to increase computation in LRTs: adding independently parameterized layers (physical depth) versus inserting latent thought tokens that reuse existing layers (temporal depth). They conduct experiments using a 2L-layer LRT and an L-layer LRT with thought tokens, measuring performance improvements and parameter efficiency.
Results
The experiments show that the LRT with one thought token achieves performance improvements close to that of a model with double the depth, while using approximately 48% fewer parameters. A second thought token further enhances performance, although models with triple depth still outperform them at the same decoding block count.
Implications
The findings suggest that incorporating temporal computation through thought tokens can lead to more efficient Transformer architectures, potentially influencing future designs of recurrent models in NLP and other domains. This approach may reduce the computational burden while maintaining or improving performance.
COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules
Theory
Optimization
- Introduces COMPLEX, a certified embedding for multiparameter persistence modules.
- Establishes the first two-sided distortion bound for multiparameter feature maps.
- Achieves state-of-the-art performance on Orbit benchmarks, surpassing existing methods.
- Demonstrates that the embedding is adaptable while maintaining certification guarantees.
Read more
COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules
Summary
This paper introduces COMPLEX, a novel closed-form embedding for multiparameter persistence modules that addresses the lack of a lower gauge in existing vectorizations. By slicing the module along a near-diagonal net and embedding each slice using the certified PLACE/PALACE landmark map, the authors achieve a two-sided distortion bound, allowing for measurable faithfulness in the features. The method demonstrates a significant improvement in performance on various benchmarks, including Orbit5k and Orbit100k, outperforming existing methods such as Euler-characteristic surfaces and transformers. The paper also discusses the implications of the embedding for local per-prediction certification and the adaptability of the method to different configurations without losing the certification guarantee. Overall, COMPLEX sets a new state of the art in multiparameter persistence feature maps, providing a robust framework for future research in this area.
Methodology
The authors propose a closed-form embedding method that slices multiparameter persistence modules along a fixed near-diagonal net. Each slice is embedded using the PLACE/PALACE landmark map, and the resulting embeddings are concatenated. The method includes a coherence condition to ensure that separated modules remain distinct in the embedding space, thus providing a measurable lower gauge.
Results
COMPLEX achieves a performance of 91.95% on the Orbit5k benchmark and 92.98% on Orbit100k, outperforming both Euler-characteristic surfaces and transformer-based methods. The method also shows improved accuracy on molecular benchmarks compared to existing multiparameter methods, clearing COX2’s majority baseline by more than three points.
Implications
The findings suggest that COMPLEX can enhance the reliability of multiparameter persistence features in machine learning applications, particularly in domains requiring certified predictions. This could lead to advancements in areas such as topological data analysis and other fields where multiparameter persistence is relevant.
OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems
NLP
Large Language Models
Graph Learning
- Introduction of OpenMAS-GCom as a benchmark for diagnosing performance in G-MAS.
- Controlled interventions allow for the isolation of the effects of communication structures and role assignments.
- Evaluation of 17 configurations across 29 datasets, including 400 complex tasks.
- Significant performance differences observed when removing specialist agents versus critic agents.
Read more
OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems
Summary
The paper introduces OpenMAS-GCom, a diagnostic benchmark designed to evaluate graph-enhanced multi-agent systems (G-MAS). These systems utilize communication graphs and role assignments to coordinate large language model agents, but existing evaluations struggle to attribute performance differences to specific organizational choices due to intertwined factors like model differences and communication patterns. OpenMAS-GCom addresses this by allowing controlled interventions on G-MAS configurations, enabling researchers to isolate the effects of communication structures, role compositions, and information flows on performance. The benchmark evaluates 17 configurations across 29 datasets in six domains, introducing 400 complex tasks that require agents to synthesize information from multiple sources. The authors demonstrate that removing specialist agents significantly impacts performance compared to critic removal, and that different configurations yield varying accuracies on complex tasks. The benchmark is open-sourced, promoting reproducibility and further research in G-MAS.
Methodology
The authors developed a benchmark that utilizes controlled organizational interventions to assess the impact of various components in G-MAS. They represent systems through collaboration units, communication links, shared information, and execution rules. The benchmark compares original systems with modified versions by changing one component at a time while keeping other factors constant. Four intervention protocols were employed: rewiring communication edges, removing specialist or critic agents, replacing messages with incorrect content, and disabling workers during execution.
Results
Experiments revealed that the removal of specialist agents led to a mean score reduction of 3.58 percentage points, compared to only 0.58 points for critic removal. Additionally, different configurations exhibited varying performance degradation under incorrect messages and worker failures, despite similar original scores. On the newly introduced G-MAS-Complex tasks, distinct configurations achieved the highest accuracy and accuracy per token.
Implications
The OpenMAS-GCom benchmark has the potential to enhance the understanding of how communication structures and role assignments affect the performance of multi-agent systems. It can guide future research in optimizing G-MAS configurations and contribute to the development of more effective collaborative AI systems.
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
Large Language Models
Reinforcement Learning
NLP
- Introduction of BI-Bench, the first benchmark for evaluating LLMs in end-to-end BI workflows.
- Demonstration of poor performance of existing LLMs on BI-Bench, highlighting the complexity of BI tasks.
- Development of BI-Agent, which decomposes BI workflows into subtasks and utilizes specialized data management methods.
- Significant accuracy improvements achieved through tool augmentation and post-training techniques.
Read more
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
Summary
This paper addresses the challenges faced by enterprise users in Business Intelligence (BI) workflows, which typically involve complex and time-consuming data preparation steps. The authors introduce BI-Bench, a novel benchmark designed to evaluate the performance of large language models (LLMs) in answering BI questions without manual data preparation. The benchmark consists of real-world BI projects and pairs of business questions with their corresponding ground-truth answers. Initial experiments reveal that existing LLMs struggle with BI-Bench, achieving less than 50% accuracy. To enhance performance, the authors propose BI-Agent, a tool-augmented system that breaks down BI workflows into manageable subtasks, such as searching, joining, and transforming data. Additionally, they develop a post-training framework that utilizes real BI project data to improve the model's performance through supervised fine-tuning and reinforcement learning. The results show significant accuracy improvements, with BI-Agent achieving up to 40 percentage points higher accuracy on BI-Bench compared to vanilla LLMs, and post-training yielding further gains. This work emphasizes the importance of integrating tool-augmented reasoning and domain-specific training in BI applications, paving the way for future research in automating BI processes.
Methodology
The authors constructed BI-Bench by crawling real-world BI projects and manually extracting business questions and their answers. They developed BI-Agent to automate BI workflows by breaking them into subtasks and orchestrating specialized methods. A post-training framework was also created to enhance the model's capabilities using supervised fine-tuning and reinforcement learning.
Results
BI-Agent demonstrated substantial accuracy improvements on BI-Bench, achieving up to 40 percentage points higher accuracy compared to vanilla LLMs. Post-training further enhanced performance, yielding gains of up to 30 points. These improvements were statistically significant, underscoring the effectiveness of the proposed methodologies.
Implications
The findings suggest that automating BI workflows using LLMs and tool-augmented reasoning can significantly ease the burden on non-technical users, making BI more accessible. This research opens avenues for further exploration in automating complex data analysis tasks and improving decision-making processes in enterprises.
What Must Survive? Exact Task-Information–State Frontiers for Resource-Sufficient Learning
Theory
Efficient ML
Optimization
- Introduces exact task-information-state frontiers for resource-efficient learning.
- Demonstrates that advance task information can reduce the necessary retained state through low-rank partitions.
- Establishes strong NP-hardness for optimal advice partitioning in linear task families.
- Provides practical examples illustrating the theoretical findings in real-world applications.
Read more
What Must Survive? Exact Task-Information–State Frontiers for Resource-Sufficient Learning
Summary
This paper investigates the relationship between task information and state retention in resource-sufficient learning systems. It addresses how much state can be saved when advance task information is limited. The author introduces the concept of exact task-information-state frontiers, which quantifies the minimal retained state necessary for exact recovery of tasks given a finite family of linear tasks. The findings reveal that advance task information allows for effective state reduction through partitions of tasks with low rank joint task operators. The paper also establishes the NP-hardness of finding optimal advice partitions and presents several examples illustrating the theoretical results, including a hierarchical multi-task model and a digital twin application. The results indicate that task-aware compression can significantly reduce the required state dimensions while maintaining performance.
Methodology
The methodology involves defining a framework where a finite family of linear tasks is considered. An oracle reveals a deterministic advice symbol before state formation, and the exact task is revealed afterward. The author formulates the problem mathematically, deriving exact and approximate frontiers based on task partitions and ranks of stacked operators. The paper employs theoretical proofs and constructs to establish the relationships between retained state dimensions and task ranks.
Results
The main results include the formulation of the exact task-information-state frontier, which is characterized by the minimum retained dimension required for exact recovery of tasks. The paper also presents a robust approximate frontier for scenarios with non-zero recovery error, demonstrating that the task partitioning problem is NP-hard. Specific examples illustrate the application of these theoretical results, showing significant reductions in required state dimensions in practical scenarios.
Implications
The findings have significant implications for designing efficient learning systems that can adapt to varying task requirements with limited resources. The theoretical framework can guide the development of algorithms for task-aware compression and predictive state representations, potentially enhancing performance in multi-task learning environments and applications requiring efficient resource management.
Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs
Theory
Generative Models
Robotics
- Gaussian uniqueness does not extend to non-Euclidean latent spaces.
- Conditions for linear recovery are derived for manifold worlds.
- Geometry-aware regularizers improve representation learning.
- Experiments show that compatible target geometries enhance recovery performance.
Read more
Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs
Summary
This paper extends the analysis of Joint-Embedding Predictive Architectures (JEPAs) beyond the traditional Gaussian assumptions, focusing on the role of latent geometry in representation learning. Previous work established that under Euclidean conditions, matching a Gaussian target distribution guarantees the recovery of latent variables up to a linear transformation. This study broadens this framework to include latent variables supported on embedded Riemannian manifolds, revealing that Gaussian uniqueness is not a universal property. The authors derive conditions under which alignment and exact distribution matching can recover latent states linearly, particularly highlighting that when latent variables are uniformly distributed on a sphere, optimal representations can recover the latent state up to an orthogonal transformation. The paper also introduces geometry-aware heat-kernel MMD regularizers for various target distributions and demonstrates through experiments that geometrically compatible targets lead to improved recovery in high-dimensional spaces. The findings emphasize the importance of aligning the representation target with the latent geometry to avoid distortion in the learned representations.
Methodology
The authors extend identifiability theory from Euclidean to non-Euclidean latent spaces using intrinsic stochastic dynamics for generating positive pairs. They implement geometry-aware heat-kernel MMD regularizers for various target distributions and conduct experiments across Gaussian, spherical, toroidal, and Clifford-torus latent spaces to validate their theoretical findings.
Results
The study finds that non-Euclidean latent geometries can support multiple linearly recoverable distributions, with the uniform sphere providing a tighter approximate-recovery guarantee than the Gaussian case. Experiments confirm that when the target geometry aligns with the latent space, the recovery of latent states is significantly improved.
Implications
These findings suggest that in practical applications of JEPAs, careful consideration of the latent geometry can enhance representation learning, particularly in complex, high-dimensional spaces. This could lead to better performance in tasks involving non-Euclidean data structures, such as robotics and computer vision.
Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources
Theory
Generative Models
Interpretability
- Proves identifiability of nICA for real analytic functions with Laplace-like source distributions.
- Introduces Real Analytic Decoders (RAD) that require minimal changes to existing machine learning frameworks.
- Demonstrates the ability to recover interpretable latent factors from complex datasets.
- Establishes a theoretical foundation for nICA that enhances training stability and reproducibility.
Read more
Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources
Summary
This paper addresses the challenge of recovering true underlying factors in nonlinear Independent Component Analysis (nICA), particularly when the source probability density functions exhibit finite discontinuities in their first derivatives. The authors prove identifiability of real analytic generating functions under these conditions, with the Laplace distribution serving as a primary example. The proof contrasts the kinks in source distributions with the smoothness of real analytic functions, establishing that nICA can achieve exact recovery of latent factors without auxiliary observations. The authors introduce a method called Real Analytic Decoders (RAD), which can be implemented with minimal modifications to existing training pipelines using Normalizing Flows or Variational Autoencoders. Extensive experiments on both synthetic and real datasets, including the CelebA dataset, demonstrate the effectiveness of RAD in recovering interpretable latent factors, thus validating the theoretical results and showcasing the method's stability in training.
Methodology
The authors utilize a theoretical approach to prove identifiability of nICA by analyzing the properties of real analytic functions and their relationship to source distributions with discontinuities. They implement their findings through experiments using Normalizing Flows and Variational Autoencoders, ensuring that the activation functions used are real analytic. The methodology includes both theoretical proofs and empirical validation on various datasets.
Results
The paper successfully demonstrates that nICA can recover true latent factors from Laplace-like sources with finite discontinuities in their first derivatives. Experiments show that RAD can effectively identify interpretable latent variables in the CelebA dataset, confirming the theoretical claims of identifiability and the practical applicability of the proposed method.
Implications
The findings have significant implications for unsupervised learning and representation learning, particularly in applications where understanding the underlying factors of complex data is crucial. The ability to recover interpretable latent representations can enhance the performance and reliability of machine learning systems in various domains, including image processing and financial analysis.
Available Guardrails: Certifying Selective Prediction across ML Systems
Theory
Optimization
Efficient ML
- Introduces the concept of certificate availability as a critical factor in selective prediction.
- Develops a dynamic programming approach to optimize reporting-partition selection.
- Implements a held-out selection procedure to maintain validity in partition certification.
- Demonstrates significant improvements in mean coverage across various machine learning applications.
Read more
Available Guardrails: Certifying Selective Prediction across ML Systems
Summary
This paper addresses the challenge of certifying selective predictions in machine learning systems, which act as safety gates by providing outputs only when predictions are deemed trustworthy. The authors highlight the importance of certificate availability, which is the ability to produce valid certificates from finite calibration data, especially when dealing with multiple reporting units. They propose a methodology that utilizes classical exact-binomial inversion to compute certificate availability and formulate the problem of reporting-partition selection as a dynamic programming task. This approach reveals a trade-off between safety, granularity, and served traffic. The authors demonstrate that a truth-informed planner can significantly improve mean coverage compared to naive estimators. They also introduce a held-out selection procedure to ensure validity is preserved during the partition selection process. The framework is evaluated across various datasets and architectures, showing its effectiveness in multiple domains including LLM tool-calling and content moderation.
Methodology
The authors employ classical exact-binomial inversion to compute certificate availability and formulate the reporting-partition selection as a dynamic programming problem. They utilize a held-out data split to construct candidate partitions and select among them, ensuring that the certification sample remains independent of the design choice.
Results
The proposed framework achieves a mean coverage improvement of 0.157 over support balancing with a truth-informed planner, compared to only 0.005 with naive estimators. Additionally, the held-out selection procedure improves mean coverage by 0.060 across various datasets and architectures, confirming the effectiveness of the approach in real-world applications.
Implications
The findings suggest that certified availability can be a valuable resource in deploying selective predictors, allowing for more reliable and granular safety certifications in machine learning systems. This has potential applications in areas such as clinical decision-making, automated moderation, and recommendation systems.
An Introduction to Compression-Based Machine Learning
Theory
Efficient ML
- Compression algorithms can be effectively utilized as machine learning methods.
- The paper introduces a design framework for compression-based machine learning.
- Empirical validation shows that compression-based methods outperform conventional baselines in malware classification.
- Design choices in compression-based methods can lead to significant accuracy improvements.
Read more
An Introduction to Compression-Based Machine Learning
Summary
This paper explores the intersection of lossless data compression algorithms and machine learning methods, proposing that any lossless compression algorithm can be transformed into a machine learning method through principles like Normalized Compression Distance (NCD) and Minimum Description Length (MDL). The authors survey existing strategies that leverage compression for machine learning tasks, emphasizing the potential of these methods in modern AI applications. They introduce a design framework for compression-based machine learning and empirically validate its effectiveness, particularly in the context of malware classification. The study reveals that compression-based methods can achieve competitive performance compared to conventional machine learning baselines, with notable accuracy improvements, especially in malware detection, where design variations can yield gains of up to 0.62 in accuracy. The paper also discusses the theoretical foundations of compression, including Kolmogorov complexity, and how these concepts can be applied to enhance machine learning models.
Methodology
The authors survey various compression-based machine learning techniques and introduce a unified design framework. They empirically validate their framework by comparing the performance of compression-based methods against conventional machine learning approaches, particularly in malware classification tasks.
Results
The study finds that compression-based machine learning methods are competitive with traditional baselines and show superior performance in malware classification. The design choices made in these methods can lead to accuracy gains of up to 0.62.
Implications
The findings suggest that leveraging compression techniques in machine learning could lead to improved models, particularly in domains like cybersecurity where traditional deep learning approaches may struggle. This work opens avenues for further research into the integration of compression and machine learning, potentially enhancing model performance across various applications.
Neural Cellular Automata Learn General Features in their Hidden Channels
Efficient ML
Computer Vision
Theory
- Neural Cellular Automata (NCAs) offer a parameter-efficient alternative to traditional deep learning models.
- The paper introduces a transfer-learning mechanism that utilizes hidden states from a teacher NCA to enhance a student NCA's performance.
- NCAs demonstrate superior generalization capabilities in few-shot learning scenarios compared to recurrent and feed-forward models.
- The hidden channels of NCAs capture general topological features, facilitating effective transfer learning.
Read more
Neural Cellular Automata Learn General Features in their Hidden Channels
Summary
This paper explores the internal dynamics of Neural Cellular Automata (NCAs) and introduces a novel transfer-learning mechanism that utilizes hidden states from a pretrained teacher model to enhance the performance of a student model in few-shot learning scenarios. The authors argue that while modern deep learning models often rely on over-parameterization, leading to overfitting, NCAs provide a parameter-efficient alternative that can generalize well with fewer parameters. The study evaluates NCAs against recurrent and feed-forward architectures on MNIST benchmarks, demonstrating that NCAs outperform these models in terms of generalization with a minimal parameter budget. The hidden channels of NCAs are shown to capture general, scale-invariant topological features, allowing for effective transfer learning. This research highlights the potential of hidden-state dynamics as a decentralized computational substrate for robust and efficient learning.
Methodology
The authors conducted experiments comparing NCAs with recurrent and feed-forward models on MNIST datasets, focusing on few-shot learning. They implemented a transfer-learning approach where hidden states from a pretrained NCA were injected into a student NCA during training to guide optimization. The models were evaluated based on their ability to generalize from limited examples, using metrics such as mean squared error and cross-entropy loss.
Results
The results indicated that NCAs outperformed the other architectures in generalization tasks, achieving strong performance with approximately 9,800 parameters. The hidden channels were found to effectively absorb morphological complexity and converge to orthogonal states, capturing scale-invariant features that facilitated few-shot learning on unseen classes.
Implications
The findings suggest that leveraging hidden-state dynamics in NCAs can lead to more robust and efficient learning models, particularly in data-limited environments. This approach may have applications in various domains requiring few-shot learning capabilities, such as image classification and pattern recognition.
SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
Large Language Models
Theory
- Introduces BENCHPROOFER, a pipeline for formal verification of LLM-generated code.
- SWE-PROOF benchmark includes 500 real-world coding tasks with formal correctness proofs.
- Demonstrates that many test-passing patches are flawed, indicating limitations of traditional testing.
- Highlights the difficulty of specification synthesis, with only 62% of generated specifications passing audits.
Read more
SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
Summary
The paper addresses the challenge of ensuring the correctness of code generated by large language models (LLMs) in software engineering. Traditional benchmarks rely on test suites that are often incomplete and can lead to memorization by the models. To overcome this, the authors introduce BENCHPROOFER, a pipeline that transforms coding tasks with known correct patches into formally verified tasks. This involves creating formal specifications for new code, summarizing existing functions with axioms, and passing through a series of mechanical and adversarial verification gates. The resulting benchmark, SWE-PROOF, consists of 500 real-world coding issues with formal correctness verification. The study reveals that a significant portion of test-passing patches are flawed, highlighting the inadequacy of structured natural language specifications. The authors find that while models struggle to write their own specifications effectively, providing a correct specification significantly improves verification outcomes. The paper emphasizes the importance of specification quality, identifying it as a key challenge in verified code generation, and suggests that faithful specification synthesis remains an open problem.
Methodology
The authors developed BENCHPROOFER, which creates formal specifications and proofs for coding tasks by summarizing existing code and using a combination of mechanical verification and adversarial audits. They evaluated the approach on two advanced LLMs, Claude Opus 4.8 and GPT-5.5, to assess the effectiveness of the generated specifications in improving verification accuracy.
Results
The evaluation found that between 25% to 50% of patches that passed tests were flawed. Providing a correct formal specification improved resolution rates from 85% to 95% for Opus 4.8 and from 81% to 94.7% for GPT-5.5. However, only 62% of synthesized specifications passed the audit, with the main failure mode being the lack of faithfulness in the specifications.
Implications
The findings suggest that formal verification can significantly enhance the reliability of code generated by LLMs, but the challenge of creating high-quality specifications must be addressed. This work could lead to improved practices in software engineering, particularly in automated code generation and verification.
SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference
Large Language Models
Efficient ML
NLP
- SpecQuant combines speculative decoding with multi-parent quantization for efficient LLM inference.
- The framework allows for dynamic routing of queries based on predicted task complexity.
- Achieves significant speed improvements (35-43%) with minimal accuracy loss (≤2%).
- Facilitates practical deployment of LLMs on consumer hardware without retraining or architecture changes.
Read more
SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference
Summary
SpecQuant introduces a novel framework for enhancing the inference efficiency of large language models (LLMs) on consumer hardware, which often face limitations due to compute and memory constraints. By integrating speculative decoding with multi-parent quantization, SpecQuant allows for adaptive inference without the need for retraining or architecture modifications. The framework generates multiple quantized model variants (INT4, FP8, FP16) from a shared base model and dynamically routes queries based on their complexity. Lightweight models are utilized for simpler tasks, while full-precision models handle more complex reasoning. This shared-weight design facilitates effective speculative decoding, minimizing compatibility issues associated with separate draft models. Evaluations on Qwen2.5 based models across MMLU, AlpacaEval, and GSM8K datasets demonstrate that SpecQuant achieves speedups of 35-43% while maintaining accuracy within a 2% degradation threshold. This advancement enables practical on-device deployment of LLMs across various hardware configurations, making it accessible for users without specialized infrastructure or expertise.
Methodology
SpecQuant employs a training-free approach that generates multiple quantized model variants from a single base model. It utilizes speculative decoding to predict token candidates using lightweight models, which are then validated by full-precision models based on the complexity of the task. The framework incorporates multi-parent quantization to create INT4, FP8, and FP16 variants, optimizing memory usage and computational efficiency.
Results
The evaluation of SpecQuant on various benchmarks (MMLU, AlpacaEval, GSM8K) shows speed improvements of 35-43% compared to traditional methods, with accuracy degradation not exceeding 2%. This demonstrates the framework's effectiveness in enhancing LLM inference without compromising performance.
Implications
SpecQuant's approach enables the deployment of large language models on consumer devices, enhancing accessibility and usability for end-users. It addresses the challenges of hardware limitations, making it feasible to run sophisticated LLMs locally without extensive computational resources or expertise.
M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection
Large Language Models
Graph Learning
Multimodal
- M2G-LLM integrates multimodal data into LLMs to enhance clinical predictions.
- The framework uses GNNs to model patient relationships and temporal dynamics.
- Enriched context vectors are injected into LLMs, allowing for joint reasoning.
- M2G-LLM outperforms existing models on clinical prediction tasks.
Read more
M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection
Summary
The paper presents M2G-LLM, a novel framework that integrates multimodal data—clinical notes, laboratory results, and medical imaging—into Large Language Models (LLMs) to enhance clinical prediction tasks. Traditional LLMs excel at processing unstructured clinical text but struggle with non-text modalities, limiting their effectiveness in healthcare. M2G-LLM addresses this by employing Graph Neural Networks (GNNs) to model temporal relationships between patient visits and propagate information across similar patients. This results in the creation of enriched multimodal context vectors that are injected into the LLM's intermediate layers, allowing for joint reasoning over both textual and non-textual data. The framework was evaluated on the MIMIC-IV and MIMIC-CXR datasets, demonstrating significant improvements in clinical prediction tasks compared to strong baseline models. The findings underscore the potential of combining LLMs' language understanding with GNNs' relational reasoning capabilities for comprehensive healthcare analysis.
Methodology
M2G-LLM constructs patient-level graphs to encode temporal and similarity relationships among patients. It learns modality-specific embeddings aligned in a shared space and integrates this information into the LLM using residual-stream conditioning, which allows for the conditioning of the LLM without altering the input prompt.
Results
M2G-LLM achieved superior performance in clinical prediction tasks, including one-year mortality prediction and 30-day readmission prediction, outperforming various baseline models. For instance, it achieved an accuracy of 78.98% and an F1 score of 64.55% for one-year mortality prediction, and an accuracy of 85.75% with an F1 score of 52.35% for 30-day readmission prediction.
Implications
The integration of multimodal data through M2G-LLM could significantly enhance clinical decision-making processes, leading to better patient outcomes. This framework may pave the way for more sophisticated healthcare analytics tools that leverage diverse data sources for improved predictive modeling.
Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
NLP
Large Language Models
Theory
- Introduces an inclined boundary for pretraining data detection that incorporates predictive entropy.
- Establishes a statistical rationale for entropy correction, enhancing member-non-member separation.
- Proposes Energy Transfer Detection (ETD) as a new method grounded in free-energy principles.
- Demonstrates significant performance improvements over existing detection methods.
Read more
Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
Summary
This paper addresses the challenge of detecting whether specific texts were part of the pretraining data for large language models (LLMs). Traditional methods rely on likelihood scores, which can misclassify predictable non-members as members due to overlapping prediction loss profiles. The authors propose a novel approach called Energy Transfer Detection (ETD), which incorporates predictive entropy to adjust prediction loss, creating an inclined boundary for better member-non-member separation. This method is grounded in a free-energy perspective, allowing for a clearer distinction between texts based on their likelihood and uncertainty. The paper presents extensive experiments demonstrating that ETD significantly outperforms existing likelihood-only methods, achieving improvements in average AUROC and TPR@5%FPR while maintaining robustness across various settings.
Methodology
The authors develop an entropy-corrected detection score that evaluates prediction loss in relation to predictive entropy. They analyze the expected membership signal and its variance, leading to the formulation of ETD, which combines prediction loss and predictive uncertainty into a residual free-energy contribution. This approach allows for a more effective separation of pretraining and non-pretraining texts.
Results
ETD achieves an average AUROC improvement of up to 3.5% and a TPR@5%FPR improvement of up to 5.1% compared to existing methods. The results indicate that the proposed method is robust across diverse experimental settings, effectively addressing the limitations of likelihood-only detection approaches.
Implications
The findings have significant implications for auditing large language models, particularly in ensuring compliance with privacy and intellectual property regulations. By improving the detection of pretraining data, this research contributes to the responsible deployment of LLMs and enhances the transparency of their training processes.
Tracing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework
Time Series
- Introduces a source-first taxonomy for zero-shot time-series forecasting based on evidence sources.
- Identifies four critical audit conditions that must be considered in zero-shot forecasting evaluations.
- Proposes a governance agenda for reporting evidence boundaries and assumptions in benchmarking.
- Highlights the importance of distinguishing between different evidence sources in zero-shot TSF systems.
Read more
Tracing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework
Summary
This paper addresses the concept of zero-shot time-series forecasting (TSF), which is typically defined as forecasting without updating model parameters for specific targets. However, the authors argue that this definition lacks clarity regarding the sources of evidence utilized by the forecasting systems. They propose a source-first taxonomy that categorizes zero-shot TSF methods based on their primary evidence sources: frozen language model prior reuse, parametric time-series pretraining, and retrieval-augmented external memory. The paper outlines four audit questions that remain relevant after identifying the evidence source: task interface, forecast object and scoring, prediction-time context, and resource budget. The authors advocate for a governance agenda that requires reporting evidence boundaries and interface assumptions alongside performance scores on zero-shot leaderboards, ensuring that benchmark progress accurately reflects transferable forecasting capabilities rather than undisclosed contextual changes.
Methodology
The authors develop a taxonomy that classifies zero-shot TSF systems by their primary evidence sources rather than by their architectural implementations. They analyze various forecasting configurations to determine the evidence boundaries and propose a checklist for minimum disclosure in benchmarking.
Results
The proposed taxonomy effectively categorizes existing zero-shot TSF methods and clarifies the distinctions between them based on their evidence sources. The audit framework provides a structured approach to evaluating and comparing these methods, emphasizing the need for transparency in reporting.
Implications
This work has significant implications for the field of time-series forecasting, as it encourages researchers and practitioners to adopt a more rigorous and transparent approach to evaluating zero-shot forecasting methods. It may lead to improved benchmarking practices and a better understanding of the capabilities and limitations of different forecasting systems.
Time series generation with spectrally aligned latent flow matching
Generative Models
Time Series
- Introduction of a spectrally-aligned latent-flow model for time series generation.
- Utilization of transform-consistency losses to preserve high-frequency spectral features.
- Theoretical guarantees for the proposed loss functions related to signal roughness.
- Empirical evidence showing improved signal quality and computational efficiency.
Read more
Time series generation with spectrally aligned latent flow matching
Summary
This paper addresses the challenges of time series generation using latent flow models, which often suffer from spectral mismatches due to latent compression. The authors propose a novel spectrally-aligned latent-flow time series generator that incorporates fine-tuning losses based on canonical signal representations, including Fourier, wavelet, and signature transforms. This approach aims to preserve the dynamical properties of the synthetic samples, ensuring that they align with the true data in terms of relevant features such as smoothness and spectral content. The methodology involves training an encoder-decoder pair with transform-consistency losses that prioritize high-frequency features, thus overcoming the limitations of traditional pointwise reconstruction losses. The authors conduct a comparative analysis against a base latent-flow model and state-of-the-art methods on real-world benchmark datasets, demonstrating the effectiveness of their approach in generating realistic time series while maintaining computational efficiency.
Methodology
The authors implemented a latent-flow model using an encoder-decoder architecture, trained with a combination of Euclidean loss and novel transform-consistency losses. These losses are designed to maintain relevant spectral features in the latent representation, allowing for fine-tuning that prioritizes high-frequency components. The model was evaluated on long-range univariate and multivariate time series datasets.
Results
The proposed spectrally-aligned latent-flow model outperformed the base model and state-of-the-art methods in terms of signal realism and computational efficiency. The quantitative results indicated that the fine-tuning with transform-consistency losses significantly improved the quality of the generated time series, as measured by discriminative scores and waveform-based metrics.
Implications
The findings suggest that incorporating spectral alignment in latent flow models can enhance the generation of synthetic time series, making them more suitable for applications such as data augmentation, imputation, forecasting, and denoising. This approach could lead to better performance in downstream tasks that rely on realistic time series data.
TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching
Large Language Models
Optimization
Multimodal
- Introduces TierKV, a framework for optimizing LLM inference on mobile devices.
- Utilizes Predictive Multi-Tier Cache Optimization (PMCO) to manage KV cache efficiently.
- Achieves up to 17.6× improvement in prefill throughput compared to existing frameworks.
- Reduces RAM-resident KV cache by 12.5–34%, allowing for longer context lengths.
Read more
TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching
Summary
The paper presents TierKV, a novel framework designed to optimize the inference of large language models (LLMs) on mobile devices by addressing the memory bottleneck associated with Key-Value (KV) caching. As LLMs are increasingly deployed on mobile platforms, the demand for long contexts in applications such as text, vision, and audio processing has led to significant challenges in memory management. Traditional methods for reducing KV cache size often result in reconstruction overhead or irreversible token loss, which can negate the benefits of memory savings. TierKV introduces Predictive Multi-Tier Cache Optimization (PMCO), which predicts future cache demands based on prefill hidden states and allocates tokens across multiple tiers (exact, low-rank, and flash-offloaded) while adhering to memory and accuracy constraints. This approach allows for full context access, eliminates the need for reactive eviction, and enables efficient runtime configuration without additional training. The framework is implemented for various mobile GPUs and demonstrates substantial improvements in throughput and memory efficiency, making it suitable for diverse workloads while maintaining accuracy.
Methodology
The authors developed a predictive resource allocation strategy that estimates future cache requirements before decoding begins. This involves a joint optimization of cache placement and compression using offline-calibrated per-layer SVD ranks. The framework employs an entropy-guided predictor to forecast cache demand based on hidden states generated during prefill, allowing for prompt-specific configurations without additional training. The implementation includes a split-path execution strategy to handle heterogeneous KV layouts efficiently.
Results
TierKV was evaluated across eight models spanning text, vision, and audio modalities on three mobile SoCs. The results showed that TierKV improved prefill throughput by up to 1.6× over llama.cpp and 17.6× over MNN-LLM. The fused kernel was 2.3× faster than naive SVD reconstruction, and the I/O scheduler achieved a 94.4% prefetch hit rate. Additionally, the framework reduced the RAM-resident KV-cache footprint by 12.5–34%, enabling a maximum supported context length increase of up to 2.6× within the same memory budget.
Implications
The advancements presented in TierKV have significant implications for the deployment of LLMs on mobile devices, particularly in applications requiring long contexts. By optimizing memory usage and computational efficiency, TierKV can enhance the performance of mobile applications in areas such as natural language processing, image and video analysis, and audio processing, while ensuring user privacy and low latency.
EnSol: an environment-aware graph neural network for molecular solubility prediction
Graph Learning
- EnSol utilizes an environment-aware approach to molecular solubility prediction, incorporating solute, solvent, and temperature interactions.
- The model employs graph neural networks to learn representations of molecular graphs and uses cross-attention for interaction-aware feature learning.
- EnSol predicts full solubility distributions, allowing for uncertainty estimation in predictions.
- The model outperformed existing solubility prediction models on benchmark datasets and demonstrated strong experimental validation.
Read more
EnSol: an environment-aware graph neural network for molecular solubility prediction
Summary
The paper introduces EnSol, a novel environment-aware probabilistic framework designed for predicting molecular solubility. Traditional models often rely on fixed-solvent assumptions and deterministic formulations, which limit their ability to accurately capture complex solute-solvent interactions and the effects of temperature. EnSol addresses these limitations by representing solute and solvent as molecular graphs and utilizing graph neural networks (GNNs) to learn separate representations for each. The model employs cross-attention mechanisms to integrate these representations, allowing for a nuanced understanding of solute-solvent interactions. Additionally, temperature is incorporated directly into the solvent representation through feature-wise modulation, enabling the model to account for temperature-dependent behaviors. EnSol predicts full solubility distributions using a mixture density network, which provides insights into experimental uncertainty. The model was evaluated on the SolProp and Leeds benchmark datasets, achieving Spearman correlations of 0.876 and 0.601, respectively, surpassing existing state-of-the-art models. Experimental validation demonstrated strong predictive performance across diverse solute-solvent pairs, with a Spearman correlation of 0.715, indicating EnSol's potential for reliable solubility prediction and solvent selection in practical applications.
Methodology
EnSol employs a graph neural network architecture to represent solute and solvent molecules as molecular graphs. It uses cross-attention to learn interaction features between solute and solvent, while temperature is integrated through feature-wise modulation. A mixture density network is utilized to predict full conditional solubility distributions, allowing for uncertainty quantification.
Results
EnSol achieved Spearman correlations of 0.876 on the SolProp dataset and 0.601 on the Leeds dataset, outperforming state-of-the-art models. In experimental validation, it maintained a Spearman correlation of 0.715 across diverse solute-solvent pairs, indicating strong predictive performance.
Implications
The development of EnSol has significant implications for accelerating molecular development processes, enhancing solvent selection, and reducing experimental bottlenecks in chemical research and industrial applications.
Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data
Federated Learning
- Introduction of FedDCN, a federated adaptation of Deep Clustering Networks.
- Utilization of synthetic data augmentations to improve robustness against non-IID data.
- Incorporation of a geometric regularization technique for latent space alignment.
- Demonstration of state-of-the-art performance on benchmark datasets.
Read more
Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data
Summary
This paper addresses the challenge of clustering high-dimensional data in a federated learning (FL) context, where data is distributed across multiple clients and privacy is a concern. Traditional deep clustering methods excel in centralized settings but struggle with non-identically independently distributed (non-IID) data, which is common in federated scenarios. The authors propose a novel framework called Federated Deep Clustering Networks (FedDCN) that generalizes existing Deep Clustering Networks (DCN) to the federated setting. FedDCN optimizes both reconstruction and clustering losses while incorporating synthetic data augmentations to enhance robustness against data heterogeneity. Additionally, a geometric regularization technique is employed to align latent spaces across clients. Experimental results demonstrate that FedDCN achieves state-of-the-art performance on benchmark datasets, effectively addressing the challenges posed by non-IID data distributions. The paper concludes by identifying future research directions in federated deep clustering.
Methodology
The methodology involves the development of FedDCN, which combines deep neural networks for representation learning with clustering mechanisms. The framework optimizes a dual loss function that includes reconstruction loss and clustering loss. To tackle the non-IID challenge, synthetic data augmentations are generated, and a geometric regularization term is introduced to ensure alignment of latent spaces across different clients.
Results
FedDCN was evaluated under both IID and non-IID conditions, showing significant improvements in clustering performance compared to existing federated deep clustering methods. The experimental results indicate that FedDCN not only maintains high-quality data partitions but also demonstrates robustness against data heterogeneity, achieving state-of-the-art results on various benchmark datasets.
Implications
The proposed FedDCN framework has significant implications for applications in privacy-sensitive domains such as healthcare, finance, and any field where data is distributed across multiple clients. It enables effective clustering of high-dimensional data without compromising data privacy, paving the way for more robust federated learning applications.
RegKT: Interpretable and Robust Deep Knowledge Tracing With IRT-Regularizer
Interpretability
- RegKT combines the strengths of IRT and DKT to enhance interpretability and robustness in knowledge tracing.
- The model incorporates a regularization term based on IRT to mitigate overfitting and improve generalization on small datasets.
- A hyper-parameter ε allows for a controlled trade-off between model accuracy and interpretability.
- The proposed approach is designed to be more applicable in real-world educational settings, addressing the needs of educators for understandable results.
Read more
RegKT: Interpretable and Robust Deep Knowledge Tracing With IRT-Regularizer
Summary
The paper presents RegKT, a novel hybrid model that integrates the interpretability of Item Response Theory (IRT) with the temporal modeling capabilities of Deep Knowledge Tracing (DKT). While deep learning models have improved the accuracy of knowledge tracing, they often lack interpretability and are prone to overfitting, particularly in educational contexts with small datasets. RegKT addresses these issues by incorporating an IRT-based regularization term into the DKT framework, allowing for a balance between accuracy and interpretability. The model's hyper-parameter ε controls this trade-off, making it suitable for real-world educational applications. The authors review related work, describe their methodology, and discuss the implications of their findings for enhancing interpretability in educational technology.
Methodology
The authors developed RegKT by integrating an IRT-based regularization term into the loss function of the DKT model. This hybrid approach allows for the modeling of student knowledge over time while maintaining interpretability through the regularization mechanism.
Results
The results demonstrate that RegKT achieves high predictive accuracy comparable to existing deep learning models while significantly improving interpretability. The model effectively reduces overfitting, making it more reliable for educational applications with limited data.
Implications
RegKT has the potential to enhance the adoption of knowledge tracing models in educational technology by providing interpretable insights into student learning processes. This could facilitate better decision-making for educators and instructional designers, ultimately leading to improved personalized learning experiences.
GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills
NLP
Large Language Models
Optimization
Graph Learning
- GraphSkillEvo introduces a graph-structured representation for agent skills, enhancing clarity and reducing redundancy.
- The framework employs evolutionary optimization techniques to explore the skill space more effectively than traditional methods.
- Extensive experiments show that GraphSkillEvo outperforms the baseline SkillOpt, improving accuracy on various benchmarks.
- The structured skill representation facilitates better execution and optimization for LLM agents, particularly for less capable models.
Read more
GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills
Summary
This paper addresses the challenges of skill optimization for Large Language Model (LLM) agents by proposing a novel representation of skills as graph-structured artifacts. Traditional skill optimization methods often rely on unstructured natural-language instructions, which can lead to inefficiencies in execution and optimization due to their lack of explicit workflow guidance and the vast search space they create. The authors introduce GraphSkillEvo, an evolutionary optimization framework that utilizes graph-structured skills, where each node represents an execution step with operational guidance, and directed edges denote context-dependent transitions. This structured representation not only clarifies the workflow but also reduces redundancy, making the optimization process more effective. The framework employs population-based evolutionary computation techniques, including mutation and crossover operators, to explore the structured skill space comprehensively. Experimental results demonstrate that GraphSkillEvo significantly outperforms existing methods, achieving notable improvements in accuracy across multiple agent benchmarks.
Methodology
The authors developed GraphSkillEvo, a population-based evolutionary optimization framework that utilizes graph-structured skills. Each skill is represented as a graph where nodes correspond to execution steps and edges represent transitions. The framework employs mutation and crossover operators to evolve a population of candidate skills, allowing for a structured search of the skill space.
Results
GraphSkillEvo consistently outperformed the baseline SkillOpt, achieving an average accuracy improvement of 4.01% on GPT-5.4-nano and 1.76% on GPT-5.4 across five agent benchmarks. The structured representation led to significant enhancements in both skill execution and optimization.
Implications
The findings suggest that graph-structured skills can greatly enhance the performance of LLM agents in various applications, enabling more effective task execution and optimization. This approach may lead to improved adaptability and efficiency in deploying LLMs across different domains, preserving procedural knowledge as models evolve.
HMB-GAN: Hybrid Multi-Bézier GAN for Vector Shape Synthesis
Generative Models
- Introduction of HMB-GAN for synthesizing CAD-ready vector geometries using multi-segment Bézier representations.
- Comparison of quantum and classical generators within the GAN framework, highlighting the advantages and limitations of each.
- Development of a Bézier decoder that enforces geometric continuity and closure in generated shapes.
- Demonstration of the feasibility of hybrid quantum-classical architectures for structured geometry modeling.
Read more
HMB-GAN: Hybrid Multi-Bézier GAN for Vector Shape Synthesis
Summary
This paper introduces HMB-GAN, a novel hybrid quantum-classical generative adversarial network designed for synthesizing CAD-ready vector geometries. Unlike previous approaches that focus on rasterized or single-Bézier representations, HMB-GAN employs a multi-segment Bézier framework to create closed shapes while ensuring geometric continuity. The architecture consists of a quantum-enhanced generator and a classical generator, evaluated against point cloud distribution metrics and geometric shape statistics. The findings reveal that while the quantum generator exhibits faster convergence and a reduced parameter count, it is hindered by excessive simulator overhead, limiting its practical application. The study emphasizes the potential of hybrid quantum architectures in modeling structured geometries while acknowledging current hardware limitations. This research contributes to the field by providing an end-to-end differentiable framework for generating complex vector shapes, thus bridging the gap in generative modeling for CAD applications.
Methodology
The HMB-GAN framework consists of three main components: a hybrid quantum generator utilizing variational quantum circuits (VQCs), a structured multi-segment Bézier decoder for outputting closed vector shapes, and a PointNet discriminator for evaluating point clouds. The generator encodes latent vectors into quantum states or higher-dimensional latents, which are then transformed into Bézier segments by the decoder. The discriminator classifies these point clouds, and gradients are backpropagated to improve the realism of generated shapes.
Results
The study found that the quantum generator achieved faster convergence and a lower parameter count compared to the classical generator. However, the quantum generator faced significant simulator overhead, which impacted its practical evaluation. The classical generator, while less efficient in convergence, was more stable under current hardware constraints. Overall, the results indicate the potential for hybrid architectures in generating structured geometries, despite existing limitations.
Implications
The findings suggest that hybrid quantum-classical GANs could significantly advance the field of computer-assisted design (CAD) by enabling the generation of complex, editable vector shapes. This could have applications in various industries, including engineering, architecture, and product design, where precise geometric representations are crucial.
Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks
Optimization
- Stiefel-AdamW optimizes linear factorization blocks by incorporating geometric considerations.
- The method stabilizes training and improves convergence rates compared to standard AdamW.
- It is validated on various models, including GPT2 and ViT, showing consistent performance gains.
- Stiefel-AdamW serves as a near drop-in replacement for AdamW, applicable across different architectures.
Read more
Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks
Summary
The paper introduces Stiefel-AdamW, a novel optimization algorithm designed for linear factorization blocks commonly found in modern deep learning architectures. These blocks, represented as W = BA, where B and A are parameter matrices, often face optimization challenges due to their non-unique factorization, which can lead to instability during training and restrict learning rates. Traditional optimization methods, such as AdamW, fail to account for the underlying geometry of these blocks, leading to potential issues. Stiefel-AdamW addresses this by constraining one factor to the Stiefel manifold while allowing the other to remain in Euclidean space. This approach maintains the coordinate-wise diagonal preconditioning of AdamW, ensuring practical efficiency while enhancing stability and convergence. The authors validate Stiefel-AdamW through experiments on LoRA-style fine-tuning of models like GPT2 and ViT, demonstrating consistent performance improvements over standard baselines without significant additional computational cost. The method is proposed as a general solution applicable to any factorization block in deep learning, making it a versatile tool for practitioners.
Methodology
The authors propose Stiefel-AdamW as a modification of the AdamW optimizer, where one factor of the linear factorization is constrained to the Stiefel manifold, while the other remains in Euclidean space. Moment estimation is performed in the ambient Euclidean space, with geometry incorporated through tangent-space projections and manifold retractions. This allows for efficient updates while preserving the benefits of coordinate-wise diagonal preconditioning.
Results
Stiefel-AdamW was tested on LoRA-style fine-tuning of GPT2, ViT, and Mistral 7B, as well as on full pretraining of GPT2 on OpenWebText. The results showed consistent improvements over strong baselines, demonstrating enhanced stability and performance without incurring significant additional computational costs compared to AdamW.
Implications
The introduction of Stiefel-AdamW has significant implications for optimizing deep learning models that utilize linear factorization blocks. Its geometry-aware approach can lead to more stable training processes and higher learning rates, making it a valuable tool for researchers and practitioners in various applications, including NLP and computer vision.
Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications
Reinforcement Learning
Robotics
Theory
- Introduction of a model-based RL method that integrates Bayesian planning for LTL specifications.
- Development of the Bayes-Adaptive Monte-Carlo Planning (BAMCP) algorithm for efficient policy synthesis.
- Demonstration of improved property satisfaction and sample efficiency over traditional model-free approaches.
- Ablation studies confirm the effectiveness of the BAMCP algorithm compared to classical methods.
Read more
Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications
Summary
This paper introduces a novel model-based Reinforcement Learning (RL) algorithm designed for efficient policy synthesis that adheres to Linear Temporal Logic (LTL) specifications in unknown environments. The authors propose a framework that synchronizes a Limit-Deterministic B¨uchi Automaton (LDBA) representation of LTL tasks with a Bayes-Adaptive Markov Decision Process (BAMDP) model of the environment. This integration allows for a more effective exploration-exploitation trade-off through Bayesian RL, contrasting with traditional non-Bayesian methods. A key contribution is the development of a Bayes-Adaptive Monte-Carlo Planning (BAMCP) algorithm, which facilitates approximate Bayes-optimal strategy synthesis within the BAMDP framework. The paper presents extensive experiments demonstrating the approach's effectiveness in achieving property satisfaction and sample efficiency compared to conventional model-free methods. Additionally, ablation studies underscore the superiority of the BAMCP algorithm over classical BAMCP in meeting LTL task requirements. The authors also highlight the application of their method in cautious RL, aiming to minimize task violations during policy training.
Methodology
The authors propose an end-to-end model-based RL approach that combines a Limit-Deterministic B¨uchi Automaton (LDBA) for LTL tasks with a Bayes-Adaptive Markov Decision Process (BAMDP) to facilitate efficient policy synthesis. The BAMCP algorithm is introduced to approximate Bayes-optimal actions for satisfying LTL objectives. The methodology involves extensive experimentation across various tasks to validate the proposed approach against existing architectures.
Results
The experimental results indicate that the proposed method significantly enhances the probability of satisfying LTL specifications while also improving sample efficiency compared to traditional model-free RL methods. The ablation studies reveal that the BAMCP algorithm outperforms classical BAMCP in achieving task satisfaction, and the model-based framework effectively reduces task violations during policy training.
Implications
This work has potential implications for developing RL agents capable of operating in complex environments with temporal constraints. The integration of Bayesian methods with LTL specifications could lead to more robust and efficient policy synthesis in safety-critical applications, such as robotics and autonomous systems.
Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications
Robotics
Audio & Speech
Efficient ML
- Proposes an alternative to image-based sonar data processing using CSV format.
- Achieves a 91.18% reduction in processing time compared to traditional methods.
- Improves object detection accuracy and enhances signal quality metrics.
- Utilizes median filtering and background subtraction for noise reduction.
Read more
Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications
Summary
This paper presents a novel approach to sonar data processing for remote sensing applications, challenging the traditional image-based methods that dominate the field. The authors propose utilizing raw sonar data in Comma-Separated Value (CSV) format, which allows for more efficient processing without the need for extensive preprocessing into image formats. This method significantly reduces computational overhead and processing time, achieving a remarkable 91.18% reduction in processing time. Additionally, the study demonstrates improvements in object detection accuracy through machine learning techniques, alongside enhancements in signal quality metrics such as Signal-to-Noise Ratio (SNR) and Peak Signal-to-Noise Ratio (PSNR). The research emphasizes the importance of optimizing data handling and processing pipelines to facilitate faster and more reliable underwater mapping and object detection, which is crucial for applications requiring real-time decision-making. The findings suggest that alternative preprocessing methods, such as median filtering and background subtraction, can effectively enhance the clarity of detected objects while preserving essential features in sonar data.
Methodology
The authors implemented an alternative data processing method that involves handling sonar data in CSV format. They applied median filtering and background subtraction techniques to enhance the quality of the data while optimizing data structures for faster interpretation. This approach minimizes the need for computationally intensive preprocessing typically associated with image-based representations.
Results
The proposed method resulted in a 91.18% reduction in processing time, improved object detection accuracy, and increased SNR and PSNR metrics. These results indicate a significant advancement in the efficiency and effectiveness of sonar data processing for remote sensing applications.
Implications
The findings of this research have significant implications for the field of underwater navigation and environmental monitoring, particularly in scenarios requiring rapid decision-making and accurate object detection. The proposed methods could enhance the capabilities of autonomous underwater vehicles and other sonar-based systems, leading to improved operational efficiency and situational awareness.
Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation
Optimization
- Introduces a single-level optimization formulation for decision-focused learning in MVO.
- Maintains budget and short-sale constraints during the learning process.
- Demonstrates improved performance on investment metrics through empirical experiments.
- Incorporates a regularization scheme to enhance model stability and prevent overfitting.
Read more
Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation
Summary
This paper addresses the limitations of traditional mean-variance portfolio optimization (MVO) frameworks, which often separate the prediction of expected returns from the optimization process, leading to misalignment between prediction accuracy and portfolio performance. The authors propose a novel decision-focused learning (DFL) approach that integrates the optimization problem directly into the learning process. By reformulating the MVO as a single-level optimization problem using the Karush-Kuhn-Tucker (KKT) optimality conditions, the proposed method maintains the budget and short-sale constraints inherent in MVO while ensuring tractability for standard optimization solvers. The authors also introduce a regularization scheme to enhance numerical stability and prevent overfitting. The effectiveness of the proposed method is evaluated through rolling-window experiments on real-world ETF data, demonstrating superior performance across multiple investment metrics compared to existing methods.
Methodology
The authors reformulate the mean-variance optimization problem as a bilevel optimization problem and then convert it into a single-level nonlinear optimization problem using the KKT optimality conditions. This approach allows for the direct minimization of decision loss while preserving the constraints of the MVO framework. A regularization scheme is also introduced to stabilize the learning process.
Results
The proposed method outperformed existing decision-focused learning approaches on multiple investment metrics in experiments conducted on real-world ETF data. The regularization scheme consistently improved performance, indicating its effectiveness in enhancing model robustness.
Implications
This work has significant implications for asset management and financial decision-making, as it provides a more effective framework for portfolio optimization that aligns predictive modeling with actual investment outcomes. The approach can be applied to various financial contexts where decision quality is critical.
When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
Large Language Models
NLP
Theory
- Introduction of the concept of scientific-judgment collapse in AI peer review.
- Demonstration of how synthetic reviews can compress rating distributions and reduce semantic diversity.
- Development of TrustReviewer, an open-source system to mitigate the effects of recursive training.
- Controlled experimental framework to study the impact of synthetic review exposure on AI reviewers.
Read more
When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
Summary
This paper investigates the recursive training of AI reviewers using large language models (LLMs) and the resulting phenomenon termed 'scientific-judgment collapse.' The authors explore how LLMs, when trained on both official and model-generated reviews, can lead to a compression of rating distributions and a reduction in semantic diversity in peer reviews. The study employs a controlled experimental framework, starting with the Llama 3.1 8B model fine-tuned on ICLR reviews from 2018 to 2023, followed by training successor models on ICLR 2024 data with varying mixtures of official and synthetic reviews. The findings reveal that increasing the proportion of synthetic reviews leads to a significant decrease in both same-paper and corpus-level semantic diversity, indicating a homogenization of judgments. To address this issue, the authors propose TrustReviewer, an open-source LLM-based system that intervenes at two stages: training-time prevention through a curated corpus and test-time correction via paired activation steering. This work highlights the risks associated with recursive AI training in scientific evaluation and offers practical solutions to maintain judgment diversity.
Methodology
The authors conducted a controlled experiment using the Llama 3.1 8B model, first fine-tuning it on official ICLR reviews from 2018-2023. They then generated synthetic reviews for ICLR 2024 papers and trained successor models on varying mixtures of official and synthetic reviews. This setup allowed them to isolate the effects of synthetic reviews on the training of AI reviewers.
Results
The study found that introducing synthetic reviews led to a compression of rating distributions and a decrease in semantic diversity, with reductions of approximately 11% for same-paper diversity and 5% for corpus-level diversity as synthetic exposure increased from 0% to 100%. This indicated a trend towards homogenization in the judgments made by AI reviewers.
Implications
The findings underscore the importance of maintaining diversity in AI-generated peer reviews to ensure robust scientific evaluation. The proposed TrustReviewer system offers a framework for mitigating the risks associated with recursive AI training, which could be beneficial for future AI-assisted scientific evaluations.
Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise
Graph Learning
- PCC refines noisy labels before GCN training without modifying the GCN architecture.
- PCC+GCN shows improved robustness across multiple conventional label-noise models.
- Achieved the best overall average rank in the NoisyGL benchmark.
- Remains competitive under instance-dependent label noise.
Read more
Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise
Summary
This paper presents a novel hybrid framework called PCC+GCN, which integrates Particle Competition and Cooperation (PCC) with Graph Convolutional Networks (GCNs) to enhance robustness against label noise in graph learning tasks. The PCC component acts as a preprocessing stage that refines noisy labels before the GCN training, identifying suspicious labeled nodes and determining whether to preserve, remove, or reassign their labels. This approach allows for improved supervision without altering the GCN architecture. The framework was evaluated on ten datasets from the NoisyGL benchmark, demonstrating significant improvements in accuracy and robustness across various conventional label-noise models, including Uniform, Pair, and Random noise, as well as instance-dependent label noise. The results indicate that PCC+GCN not only outperforms baseline GCNs but also maintains competitive performance with other robust methods while being the fastest on eight of the ten datasets. This work highlights the effectiveness of PCC-based label refinement as a computationally efficient strategy for enhancing GCN performance in the presence of noisy supervision.
Methodology
The proposed PCC+GCN framework utilizes Particle Competition and Cooperation for label refinement, identifying and correcting noisy labels before training the GCN. The PCC component employs particle dynamics to assess label reliability, while the GCN is trained on the original graph structure and node features. The methodology includes extensive hyperparameter analysis and empirical comparisons against established robust graph learning methods under various noise conditions.
Results
PCC+GCN achieved the highest overall average accuracy and the best average rank among evaluated methods in the NoisyGL benchmark, with an average accuracy gain of 1.67 percentage points over baseline GCNs. Under instance-dependent noise, it remained competitive with leading robust methods while being the fastest on eight datasets, indicating its efficiency and effectiveness.
Implications
The findings suggest that integrating PCC with GCNs can significantly enhance the robustness of graph learning models in real-world applications where label noise is prevalent. This approach could be particularly beneficial in domains such as social network analysis, recommendation systems, and biological network modeling, where accurate label information is crucial for performance.
On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation
Theory
Interpretability
- MCR2 can lead to complete prediction failures under distribution shifts.
- Optimal coding does not guarantee reliable predictions in new environments.
- Incorporating invariance principles from IRM and REx does not eliminate OOD failures.
- The study provides a systematic analysis of MCR2's limitations for OOD generalization.
Read more
On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation
Summary
This paper investigates the limitations of the Maximal Coding Rate Reduction (MCR2) framework in the context of out-of-distribution (OOD) generalization. While MCR2 has been recognized for its potential to create interpretable and structured representations in deep learning, the authors demonstrate that it can lead to significant prediction failures when faced with distribution shifts. They identify two primary limitations: first, that MCR2 can yield complete prediction failure even when optimal coding is achieved, particularly when relying on unstable environmental features. Second, the incorporation of invariance principles from successful OOD learning methods, such as Invariant Risk Minimization (IRM) and Risk Extrapolation (REx), does not resolve these failures. The authors provide a systematic analysis of these issues, illustrating that optimal coding geometry does not guarantee reliable predictions in new environments. Their findings highlight the need for additional assumptions or learning principles to ensure stable predictive relationships across varying environments, thus raising concerns about the reliability of MCR2 in practical applications.
Methodology
The authors analyze the OOD generalization capabilities of MCR2 through a simple two-class model, isolating failure mechanisms and demonstrating the limitations of the framework in both noiseless and noisy scenarios. They also explore the effects of incorporating invariance principles into the MCR2 framework.
Results
The analysis reveals that MCR2 can achieve optimal coding while still resulting in high prediction errors, particularly when relying on unstable environmental features. Even when all possible test inputs are present during training, the coding quality can be optimal while prediction accuracy approaches zero. The incorporation of invariance principles does not mitigate these issues.
Implications
The findings suggest that while MCR2 has theoretical advantages, its practical application in real-world scenarios may be limited due to its susceptibility to distribution shifts. This highlights the need for further research into robust learning principles that can ensure reliable predictions across varying environments.
RACER: Role-Aligned Competence Estimation for Human-AI Routing
Theory
- RACER combines query-dependent adaptation with identity-free competence estimation for unseen experts.
- The framework utilizes role-relative information to improve routing decisions in human-AI collaboration.
- Empirical results show RACER's superior performance on synthetic benchmarks and real-world datasets.
- The method maintains calibration across various metrics, enhancing the reliability of deferral decisions.
Read more
RACER: Role-Aligned Competence Estimation for Human-AI Routing
Summary
The paper introduces RACER, a novel framework for estimating the competence of unseen experts in human-AI collaboration settings, particularly in high-stakes prediction tasks. Traditional learning to defer (L2D) methods struggle with adapting to new experts and maintaining performance across varying contexts. RACER addresses these challenges by providing a role-relative competence estimation that is both query-dependent and invariant to class relabeling. The framework utilizes a posterior-predictive probability to assess an expert's correctness for a given query and candidate role, integrating this with model predictions to optimize decision-making. The authors demonstrate that RACER outperforms existing methods on synthetic benchmarks and real-world datasets, including chest radiography tasks, by effectively leveraging additional context and maintaining calibration across different metrics. The results indicate that RACER can enhance human-AI collaboration by improving the accuracy of deferral decisions, ultimately leading to better outcomes in critical applications such as healthcare.
Methodology
RACER employs a role-aligned competence estimation approach that uses nonparametric and neural kernel-pooling methods to evaluate an unseen expert's correctness based on context. It estimates the posterior-predictive probability of expert correctness for each candidate class role and integrates this with the model's predictions to derive a Bayes-relevant correctness probability. The framework is designed to be invariant to class relabeling while allowing for instance-dependent adaptation.
Results
RACER demonstrated strong performance on controlled synthetic benchmarks, particularly in a PathMNIST histopathology context-scaling study. It achieved the best aggregate performance on unseen expert splits in CIFAR-100 synthetic experiments. In real-world applications, such as the VinDr-CXR and CheXpert chest radiography benchmarks, RACER was competitive or the best in budget-swept deferral, with varying calibration results across different metrics and datasets.
Implications
The findings suggest that RACER can significantly improve human-AI collaboration in critical decision-making scenarios, such as healthcare, by providing more accurate and reliable deferral decisions. This could lead to better patient outcomes and more efficient use of expert resources.
MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery
Interpretability
Optimization
Theory
- MOSAIC-SR combines neural generation with symbolic search for improved equation recovery.
- The framework utilizes a pretrained Transformer to propose initial equation sketches.
- Local search and scale-aware constant fitting are employed to refine equations post-initialization.
- MOSAIC-SR outperforms existing methods in both symbolic solution rates and predictive accuracy.
Read more
MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery
Summary
MOSAIC-SR introduces a novel approach to symbolic regression, which aims to recover closed-form equations from observational data. Traditional methods face challenges in balancing flexible structural search with efficient inference, often relying on costly combinatorial optimization or generating formulas that may contain symbolic errors. MOSAIC-SR leverages a pretrained Transformer model to propose multiple initial sketches of equations, which guide the search process in promising regions of the expression space, thereby avoiding random initializations. The methodology combines a Monte Carlo tree search (MCTS) to map variable tokens to observed inputs and produces complete equations, followed by a local search that refines both the structure and constants of the equations. This approach emphasizes the importance of numerical optimization and symbolic repair in achieving high predictive accuracy and correct symbolic recovery. The framework was evaluated on the SRSD-Feynman dataset and six additional benchmarks, demonstrating superior performance in symbolic solution rates and predictive accuracy compared to existing methods. The results indicate that learned priors can effectively focus the search process, enhancing the recovery of scientific equations.
Methodology
MOSAIC-SR employs a hybrid framework that integrates a pretrained Transformer for generating initial symbolic sketches and a Monte Carlo tree search (MCTS) for mapping these sketches to observed data. The framework then conducts a local search that alternates between structural edits and constant refitting, allowing for precise symbolic recovery and high predictive accuracy.
Results
MOSAIC-SR achieved the highest symbolic solution rate across all tested datasets, including the SRSD-Feynman dataset, while also ranking among the top two methods in predictive accuracy. The framework maintained its performance even in the presence of irrelevant dummy variables, demonstrating robustness and effectiveness in scientific equation recovery.
Implications
The findings suggest that MOSAIC-SR can significantly enhance the process of scientific discovery by providing interpretable models that accurately represent underlying physical laws. This approach could be applied in various scientific fields where understanding the relationship between variables is crucial.
Sparse Priors for Efficient Distribution Learning
Theory
Generative Models
Efficient ML
- Introduction of sparse priors and the concept of Sparse Dimension to measure prior sparsity.
- Demonstration of a Bayesian risk lower bound of Ω(√k/n) for k-sparse priors, improving upon traditional bounds.
- Establishment of statistical equivalence between distribution learning and learning to sample.
- Proposed priors create well-separated clusters, enhancing the learning efficiency in high-dimensional spaces.
Read more
Sparse Priors for Efficient Distribution Learning
Summary
This paper addresses the limitations of existing theoretical guarantees in distribution learning, particularly the curse of dimensionality that affects sample complexity. The authors introduce the concept of 'sparse priors' and define a new measure called 'Sparse Dimension', which quantifies the sparsity of a prior over the space of distributions. They demonstrate that using a k-sparse prior allows for a Bayesian risk lower bound of Ω(√k/n) under common distance metrics, such as total variation and Wasserstein distances. This represents a significant improvement over traditional worst-case bounds, which degrade as O(n−1/Θ(d)). The authors argue that existing smoothness assumptions are insufficient for capturing the structure of real-world distributions, and propose a more realistic prior that creates well-separated clusters of distributions. They establish the statistical equivalence between distribution learning and learning to sample in the Bayesian framework, allowing their results to apply broadly. The findings suggest that with appropriate priors, the curse of dimensionality can be mitigated, leading to more efficient learning from fewer samples. Overall, this work provides a theoretical foundation for understanding the performance of generative models in practice, bridging the gap between theoretical predictions and empirical observations.
Methodology
The authors develop a theoretical framework that defines sparse priors and Sparse Dimension. They analyze the Bayesian risk associated with distribution learning under these priors and derive lower and upper bounds for the risk using total variation and Wasserstein distances. The paper also includes a comparative analysis of existing methods and their limitations, leading to the introduction of their proposed approach.
Results
The paper establishes that a k-sparse prior achieves a Bayesian risk lower bound of Ω(√k/n) and provides a distribution estimator that matches this bound up to logarithmic terms. The results show that learning under sparse priors can significantly improve sample efficiency and mitigate the curse of dimensionality.
Implications
The findings have significant implications for the design of generative models and distribution learning algorithms, suggesting that incorporating sparse priors can lead to more efficient learning in high-dimensional settings. This could enhance applications in various fields, including image generation, natural language processing, and other areas where generative models are employed.
OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios
Reinforcement Learning
Generative Models
Optimization
- OneBid unifies multiple oCPX advertising scenarios into a single model, enhancing efficiency and performance.
- The model employs a sequence-level Mixture-of-Experts architecture to manage cross-scenario knowledge and scenario-specific dynamics.
- CROP, a novel offline policy optimization method, mitigates risks associated with online exploration during model deployment.
- OneBid has been successfully deployed at Kuaishou, showing significant improvements in advertising performance metrics.
Read more
OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios
Summary
The paper introduces OneBid, a unified auto-bidding foundation model designed to optimize advertising strategies across diverse cost-per-X (oCPX) scenarios. Traditional auto-bidding methods have evolved from rule-based systems to reinforcement learning and generative models, yet they often operate in isolation for different advertising scenarios, leading to inefficiencies and missed opportunities for cross-scenario learning. OneBid addresses this fragmentation by integrating multiple oCPX scenarios into a single model, overcoming challenges such as multi-objective control, scalability under latency constraints, and safe offline policy improvement. The model employs a sequence-level Mixture-of-Experts architecture that balances shared knowledge across scenarios with specialized expertise for individual scenarios. Additionally, it utilizes a novel Critic-guided Relative Offline Policy optimization method (CROP) to refine its bidding strategies without the risks associated with online exploration. The effectiveness of OneBid is validated through rigorous online A/B testing, demonstrating significant performance improvements in conversion scenarios, particularly a +2.2% increase in Average Daily Value per Visitor (ADVV) and a peak of +13.1% in Return on Advertising Spend (ROAS) for specific scenarios.
Methodology
OneBid utilizes a backbone model based on Decision Transformer (DT) architecture, extending its capabilities with two atomic signals—Return-to-Go (RTG) for conversion value and Cost-to-Go (CTG) for cost efficiency. The model is pre-trained on heterogeneous oCPX logs and fine-tuned through offline post-training using the CROP method, which evaluates candidate actions in a group-relative manner to ensure safety and effectiveness.
Results
The deployment of OneBid resulted in a +2.2% improvement in Average Daily Value per Visitor (ADVV) across various conversion scenarios, with a peak improvement of +13.1% in the Return on Advertising Spend (ROAS) scenario, demonstrating the model's effectiveness in real-world applications.
Implications
OneBid's approach to unifying auto-bidding strategies can significantly streamline advertising operations, reduce engineering overhead, and enhance the ability to adapt to diverse advertising scenarios. This model sets a precedent for future research in computational advertising and the application of foundation models in other domains.
A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters
Time Series
- Introduces a lightweight pre-encoder gate for regulating covariate representations in Transformer-based forecasting models.
- Demonstrates the effectiveness of a usage-regularized variant to control average admission without redesigning the forecasting backbone.
- Evaluates the proposed method across multiple datasets and Transformer architectures, showing competitive performance.
- Explores the impact of gate placement and initialization on forecasting accuracy.
Read more
A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters
Summary
This paper addresses the challenge of integrating external covariates into Transformer-based time-series forecasting models. Traditional approaches often lack an explicit mechanism for admitting covariate information into the forecasting path. The authors propose a lightweight plug-in gate that acts as a pre-encoder admission interface, allowing for the regulation of covariate representations before they enter the encoder. This gate assigns sigmoid scores to representation units, effectively controlling their influence on the forecasting process. Additionally, a usage-regularized variant is introduced to penalize excessive admission of covariates, aiming to maintain forecasting accuracy while reducing average admission scores. The proposed method is evaluated on various Transformer architectures, including TimeXer, iTransformer, and Patch Time Series Transformer (PatchTST), under a zero-extra-tuning protocol. The experiments utilize multiple datasets, including ETTm1, ETTm2, Traffic, Energy, and influenza-like illness (ILI), and involve comprehensive analyses such as paired forecasting comparisons, gate-placement ablation, and variance inflation factor (VIF)-informed permutation feature importance (PFI) diagnostics. The results indicate that the plug-in gate is competitive with baseline models, demonstrating its effectiveness in enhancing forecasting performance while managing covariate admission.
Methodology
The authors implemented a lightweight representation-level pre-encoder gate that computes scores for covariate representation units before they enter the encoder. This gate uses a two-layer multilayer perceptron (MLP) scoring rule and is evaluated as a plug-in module across various Transformer architectures. The study includes paired forecasting comparisons, ablation studies, and controlled analyses of covariate admission, alongside a diagnostic case study using VIF and PFI.
Results
The experiments reveal that the plug-in gate maintains or improves forecasting accuracy compared to baseline models. The usage-regularized variant successfully reduces average admission scores while keeping forecasting errors comparable to the unpenalized settings. The placement of the gate and its sensitivity to initialization are also examined, providing insights into its operational effectiveness.
Implications
The findings suggest that incorporating a pre-encoder admission mechanism can enhance the performance of Transformer-based time-series forecasting models, particularly in scenarios with rich covariate information. This approach could be beneficial in various applications, including energy management, transportation planning, and public health analysis, where accurate long-term forecasting is critical.