AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

70 Papers today
8h Update frequency
7 Days of history
Benchmarking World Models for Continual Learning on Compositional Tasks
Haoyu Zhou, Joe Watson, Anson Lei, Ingmar Posner
Reinforcement Learning Robotics
  • Introduction of a compositional continual learning benchmark for world models in robot manipulation.
  • Separation of world models into task-agnostic backbones and task-specific heads to enhance continual learning.
  • Evaluation of both monolithic and modular world models to assess their performance in compositional tasks.
  • Modular world models demonstrate improved knowledge reuse but do not completely solve the forgetting problem.
Read more
Common Cause, Not Cross-Attention: Blocking Visual Shortcuts in Audio-Video Generation
Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
Generative Models Multimodal Theory
  • Models that use cross-attention or shared latents often rely on visual shortcuts, leading to incorrect audio predictions when appearance-event correlations are broken.
  • Simply sharing a latent representation does not eliminate the visual shortcut problem.
  • Counterfactual invariance is necessary and sufficient to identify the causal predictor in AV generation.
  • Interventions on nuisance variables are required to block visual shortcuts effectively.
Read more
Overlay_dx - Automating forecasting evaluation
Long H. Ngo, Mohammed Amine Chamli, Jonathan Rivalan, Thomas Jaillon
Time Series Optimization Interpretability
  • Introduction of overlay_dx as a visual evaluation metric for time series predictions.
  • Combines visual interpretability with quantitative scoring through area under the overlay curve.
  • Demonstrates effectiveness across various time series prediction scenarios.
  • Enhances communication of model performance to non-technical stakeholders.
Read more
SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration
Linhan Luo, Lequan Lin, Dai Shi, Feng Chen, José Miguel Hernández-Lobato, Junbin Gao
NLP Large Language Models
  • Introduces SUPPORTCAL, a label-free calibration method for PoLMs using PLM references.
  • Demonstrates that moderate inclusion of disagreement examples can improve calibration.
  • Retains unit weight for agreement examples while assigning continuous weights to disagreements.
  • Empirical results show lower Expected Calibration Error (ECE) compared to baseline methods.
Read more
K-TRAIL: Simulator-Guided Generative Design of EM/RF Circuits
Piyush Saha, Evan Newell, Hanna O'Leary, Arun Natarajan, Alireza Aghasi
Generative Models Optimization
  • K-TRAIL combines diffusion models with Kalman correction for RF circuit design.
  • The framework allows for automated synthesis from both S-parameter targets and RF constraints.
  • Simulator-guided generation improves layout response accuracy and diversity.
  • K-TRAIL addresses the non-uniqueness of inverse EM design effectively.
Read more
Knowledge-Graph-Augmented Chronos-2 for HEC-RAS Surrogate Forecasting
Edward Holmberg, Elias Ioup, Mahdi Abdelguerfi
Time Series Graph Learning Optimization
  • KG-Chronos-2 integrates a frozen time-series model with hydraulic project knowledge for improved forecasting.
  • The method outperformed existing models, achieving the lowest RMSE among the evaluated systems.
  • The study demonstrates the effectiveness of knowledge graphs in enhancing surrogate modeling for hydraulic forecasting.
  • Results indicate significant reductions in forecasting errors compared to traditional models.
Read more
Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise
Fabricio Breve
Graph Learning
  • PCC refines noisy labels before GCN training without modifying the GCN architecture.
  • PCC+GCN shows improved robustness across multiple conventional label-noise models.
  • Achieved the best overall average rank in the NoisyGL benchmark.
  • Remains competitive under instance-dependent label noise.
Read more
Efficient Mixture-of-Experts with Speculative Decoding via Expert Coactivation
Kumari Nishu, Han-Byul Kim, Santosh Chilkunda, Maxwell Horton, Arnav Kundu, Mohammad Samragh, Lauren Hannah, Mohammad Sekhavat, Nikhil Bhendawade, Manuel Ciosici, Iman Mirzadeh, Keivan Alizadeh Vahid, David Harrison, Irina Belousova, Mehrdad Farajtabar, Minsik Cho
NLP Large Language Models Efficient ML
  • Combining MoEs with SD can lead to significant inference speed improvements.
  • High expert coactivation during training enhances runtime efficiency.
  • Four specific training modifications can improve expert coactivation.
  • The proposed method achieves a 21% increase in throughput over standard MoEs.
Read more
Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources
Isaac Manring, Kejun Huang
Theory Generative Models Interpretability
  • Proves identifiability of nICA for real analytic functions with Laplace-like sources.
  • Introduces Real Analytic Decoders (RAD) that can be easily integrated into existing frameworks.
  • Demonstrates the method's effectiveness through experiments on synthetic and real datasets.
  • Highlights the importance of smoothness in functions for achieving identifiability in nICA.
Read more
Blind Thermodynamic Ontology Discovery from Anonymous Experiments
Linzhe Zhang, Changming Xu
Theory
  • Establishes an operational foundation for thermodynamic ontology discovery from anonymous experiments.
  • Develops a polynomial-time algorithm that extracts extensive and intensive scaling sectors.
  • Validates the framework through blind evaluations on simulated and real fluid datasets.
  • Demonstrates that omitting any of the four experimental operations leads to unresolved physical ambiguities.
Read more
Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents
Param Raval, Rohit Shenoy, Archana Vaidheeswaran
NLP Large Language Models Interpretability
  • LLM explainers can produce fluent but incorrect narratives, undermining operator oversight.
  • Significant interpretability failures were observed, including belief-drift blindness and sycophantic rationalization.
  • The study highlights the need for auditing LLM explainers in agentic deployments.
  • Proposed mitigations for identified failures were not evaluated, indicating a gap for future research.
Read more
A Distributional Optimisation Perspective on Combining Models in Deep Learning
Congye Wang, Yan Lin, Zheyang Shen, Matthew A. Fisher, Chris. J. Oates
Optimization Theory Large Language Models
  • Distributional optimisation provides a unified framework for combining models in deep learning.
  • The paper reformulates ensemble methods and LoRA averaging as entropy-regularised optimisation problems.
  • The ensemble method is convex, allowing for stronger convergence guarantees compared to LoRA averaging.
  • Empirical studies demonstrate the effectiveness of the proposed methods on various tasks.
Read more
What Must Survive? Exact Task-Information--State Frontiers for Resource-Sufficient Learning
Ronald Katende
Theory Efficient ML Optimization
  • Introduces exact task-information-state frontiers for resource-efficient learning.
  • Demonstrates how advance task information can reduce the required state dimension.
  • Establishes NP-hardness of finding optimal advice partitions.
  • Presents theoretical results supported by practical examples.
Read more
Discovering Physical Representation Languages
Linzhe Zhang, Changming Xu
Theory
  • Introduces the concept of physical representation-language discovery from controlled experiments.
  • Develops a polynomial-time procedure for recovering measurement types and their relationships.
  • Establishes an end-to-end measurement bound and matching minimax rates for finite experiments.
  • Demonstrates robustness of the framework across various physical conditions and complexities.
Read more
Multiple latent orderings better predict language model preferences
Aviral Chawla, William H.W. Thompson, Jean-Gabriel Young
NLP Large Language Models Theory
  • Intransitivity in LLM preferences indicates multiple latent orderings rather than sampling noise.
  • The proposed MBT model outperforms traditional single-utility models in predicting LLM preferences.
  • Aggregate evaluations can obscure the underlying preference heterogeneity among models.
  • Plural preferences in LLMs necessitate a reevaluation of alignment and evaluation methodologies.
Read more
Computationally efficient safe exploration in reinforcement learning
Shreeram Murali, Shankar A. Deka, Dominik Baumann
Reinforcement Learning Robotics Efficient ML
  • Introduction of COLSAFE-MDP for safe exploration in constrained MDPs.
  • Utilization of the Nadaraya-Watson estimator for constant-time updates.
  • Theoretical guarantees of safety and near-optimality with mild assumptions.
  • Significant performance improvements over GP-based methods in terms of safety and computational efficiency.
Read more
On Emergent Capabilities and Model Merging
Luca Zhou, Emanuele Rodolà
Theory Interpretability
  • Merging two models with shared emergent capabilities preserves those capabilities.
  • Emergent capabilities that are superadditive cannot be recreated through merging.
  • When only one model possesses an emergent capability, merging dilutes it faster than trained capabilities.
  • The behavior of emergent capabilities under merging differs significantly from that of trained capabilities.
Read more
TTSE: A Two-Track Online Self-Evolution Framework
Ruimin Pei, Yongkang Wu, Shangyi Zheng, Yaqing Zhang, Deyang Li, Jianjun Tao, Xinyu Zhang, Xiang Zhang
Large Language Models Reinforcement Learning Theory
  • TTSE introduces a dual-track evolution mechanism for LLM agents, separating environmental knowledge and execution strategies.
  • The framework allows continuous verification and adaptation of environmental facts and task execution procedures.
  • A decision-theoretic analysis provides insights into the conditions under which environment-conditioned policies outperform others.
  • TTSE shows improved adaptability and performance on standard benchmarks compared to existing methods.
Read more
Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization
Yifei Sheng, Haoxiang Ren, Zhilong Zhang, Haonan Wang, Runjie Xu, Yihao Sun, Nan Tang, Zhichao Wu, Lei Yuan, Haoxin Lin, Yang Yu
Reinforcement Learning Robotics Optimization
  • Introduces U-GROW, a method that prioritizes high-uncertainty states for policy optimization in VLA models.
  • Demonstrates that policy uncertainty can effectively identify states with greater potential for improvement.
  • Shows that U-GROW can be integrated into existing MBRL pipelines without changing optimization objectives.
  • Achieves improved sample efficiency and success rates in both simulated and real-world tasks.
Read more
A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters
Hongkai Zhuang, Tao Huang, Chen Hou
Time Series
  • Introduction of a lightweight pre-encoder gate for regulating covariate admission in Transformer-based forecasting models.
  • Implementation of a usage-regularized variant to control average admission without redesigning the forecasting backbone.
  • Comprehensive evaluation across multiple datasets and Transformer architectures, demonstrating the effectiveness of the proposed method.
Read more
$λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource
Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa
Reinforcement Learning Generative Models Optimization
  • Introduces path variance as a key factor in the instability of reinforcement learning for flow-matching models.
  • Proposes λ-Controlled GRPO, which uses analytical calibration instead of empirical stabilizers.
  • Demonstrates significant improvements in text accuracy and preference rewards in image generation tasks.
  • Establishes a method for budgeting gradient effort based on predicted path variance costs.
Read more
On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation
Menghui Zhou, Gaoshan Bi, Vitaveska Lanfranchi, Po Yang
Theory Interpretability
  • MCR2 can lead to complete prediction failures under distribution shifts.
  • Representations based on unstable environmental features may achieve optimal coding rates but fail in prediction accuracy.
  • Incorporating invariance principles from IRM does not eliminate OOD prediction failures.
  • The study highlights the need for new assumptions or learning principles for reliable OOD generalization.
Read more
Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning
Hao Yang, Zhenguo Xu
Reinforcement Learning Time Series Optimization
  • Extension of adversarial reinforcement learning to include Hawkes order arrivals and price impact.
  • Introduction of LSTM to handle increased non-stationarity in market conditions.
  • Characterization of equilibrium properties from both theoretical and empirical perspectives.
  • Development of a robustness evaluation protocol focusing on left-tail return metrics.
Read more
Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks
Avinash Amballa, Yashas Malur Saidutta, Wenbo Li, Lazar Valkov, Srinivas Chappidi
Efficient ML Computer Vision NLP
  • Identifies uniform rank allocation as a major source of performance gaps in LoRA merging.
  • Introduces Net Utility, a data-free metric for optimal rank allocation across tasks.
  • Demonstrates significant performance improvements in multi-task learning with non-uniform rank allocation.
  • Tests the method across diverse vision and language tasks, showing consistent enhancements.
Read more
Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks
Adewumi Augustine Adepitan, Christopher J. Haruna, Oluwasegun Adegoke, Ayooluwatomiwa Ajiboye, Oluwatobi Oluwasakin
Reinforcement Learning Optimization
  • Introduces a shared latent-space framework for urban transportation calibration and control.
  • Develops a combinatorial MLP-autoencoder for efficient simulator calibration.
  • Implements a deep Q-learning agent for dynamic traffic optimization.
  • Achieves up to 51% reduction in system-wide travel times in empirical evaluations.
Read more
When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
Sy-Tuyen Ho, Minghui Liu, Furong Huang
Large Language Models NLP Theory
  • Identification of 'scientific-judgment collapse' in AI reviewer training.
  • Controlled experiments demonstrate the impact of synthetic reviews on judgment diversity.
  • Introduction of TrustReviewer, an LLM-based system to mitigate collapse effects.
  • Findings suggest that increased synthetic review exposure compresses rating distributions.
Read more
From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning
Majid Lotfian Delouee, Hamed Ayoobi, Sjors G. J. G. In 't Veld, Martijn C. Schut
Interpretability
  • Introduces the concept of latent biomarkers for tabular clinical data.
  • Develops a combined attribution method for translating embedding-space rules into raw clinical features.
  • Demonstrates improved performance of translated rules over raw-feature rules in five out of six clinical datasets.
  • Highlights the importance of interpretable machine learning in clinical decision-making.
Read more
Explainable Neuro-Fuzzy Prediction for Trustworthy Decision-Making in Maritime
Dionisis Kalogeropoulos, Georgia Sovatzidi, Dimitris K. Iakovidis
Interpretability
  • Introduces a novel explainable decision-making framework for predictive maintenance in maritime systems.
  • Integrates a neuro-fuzzy prediction model with a two-stage explainable component.
  • Achieves an average AUC-ROC value of up to 99% on benchmark datasets.
  • Provides both feature-level and local rule-based explanations for predictions.
Read more
Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data
Morris Stallmann, Charalampos S. Kouzinopoulos, Marcin Pietrasik, Anna Wilbik
Federated Learning
  • FedDCN is the first to generalize DCN to the federated setting.
  • It operates without a centralized dataset, enhancing privacy.
  • The framework employs synthetic data augmentation and geometric regularization for improved robustness.
  • Experimental results show superior performance on benchmark datasets under IID and non-IID conditions.
Read more
MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery
Peiyi Zheng, Yanming Kang, Hans De Sterck, Giang Tran
Interpretability Optimization Theory
  • MOSAIC-SR combines neural generation with symbolic search for improved equation recovery.
  • The method uses a pretrained Transformer to propose initial equation sketches, avoiding random initialization.
  • It achieves the highest symbolic solution rates across multiple datasets while maintaining high predictive accuracy.
  • The framework emphasizes the importance of structural repair and constant fitting for accurate symbolic recovery.
Read more
GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning
Alvaro Serra-Gomez, Thomas Moerland
Reinforcement Learning Robotics Optimization
  • GEM-MPC improves the alignment between planning and learning in reinforcement learning.
  • The method combines a greedy planner-cloning policy with an exploratory KL-regularized policy.
  • Gated Prior Distillation filters stale planning targets, reducing computational costs.
  • GEM-MPC consistently outperforms existing baselines in continuous-control tasks.
Read more
ARM: Attention with Routed-Memory for Learnable Sparse Control
Qiuhao Zeng, Jerry Huang, Peng Lu, Ruiyi Fang, Gezheng Xu, Zihao Jing, Yufei Cui, Charles Ling, Gang Niu, Boyu Wang
NLP Large Language Models Efficient ML
  • ARM introduces a learnable soft eviction policy for KV-cache management, enhancing memory preservation.
  • The method employs a dynamic retrieval policy that adapts to input complexity, improving attention sparsity.
  • Experimental results show ARM's superior performance and efficiency compared to traditional KV-caching methods.
  • ARM's design allows for scalable long-context inference, addressing the growing memory demands of LLMs.
Read more
Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping
Huikang Liu, Zhengchao Wang, Daniel Kuhn, Wolfram Wiesemann
Theory Optimization
  • Introduces an index policy for MAB derived from SAB minimax solutions.
  • Establishes a semi-infinite linear programming formulation for optimizing stopping policies.
  • Achieves a distribution-free regret bound matching the minimax-optimal order.
  • Demonstrates superior performance of the proposed policy in numerical experiments.
Read more
Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision
Florian Krone, Elena Hoemann, Sven Hallerbach
Computer Vision Reinforcement Learning Optimization
  • RIBA formulates adversarial perturbation generation as a Markov decision process, enabling the use of reinforcement learning techniques.
  • The proposed method significantly reduces the number of queries needed to generate adversarial examples compared to existing black-box attacks.
  • RIBA achieves performance comparable to white-box attacks on adversarially trained models, demonstrating its effectiveness.
  • The approach is particularly relevant for real-world applications where query costs are a concern, such as in web APIs.
Read more
Lifted Bellman Linear Programming for Offline Reinforcement Learning
Hyukjun Yang, Jongchan Park, Narim Jeong, Donghwan Lee
Reinforcement Learning Optimization Theory
  • Introduction of the Lifted Bellman Linear Program (LBLP) for offline RL, focusing on in-sample Bellman optimality.
  • LBLP avoids regression losses and target networks, stabilizing the training process.
  • The unique minimizer of LBLP provides bounds between the best dataset return and the optimal value.
  • Approximate Lifted Bellman Unconstrained Minimization (ALBUM) implements LBLP with neural networks, detaching K-step rollout targets.
Read more
DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction
Ge Kong
Multimodal
  • DPTM-DT integrates multiple representation modalities for improved drug-target prediction.
  • The model employs a dual-pretrained approach with GROVER, ESM, and CTD features.
  • It utilizes a shared representation for multitask learning across regression and classification tasks.
  • Experimental results show superior performance compared to existing methods on benchmark datasets.
Read more
Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability
Vincent Souveton
Generative Models Theory Interpretability
  • Introduction of Riemannian Neural Hamiltonian Flows (RNHF) for generative modeling on Riemannian manifolds.
  • RNHF combines fixed kinetic energy from Riemannian metrics with a learned scalar potential and a geodesic integrator.
  • The model is symplectic, reversible, and volume-preserving, addressing the challenges of applying generative models to curved spaces.
  • A framework for interpretability is established, linking learned potentials to implicit profiles and target distributions.
Read more
ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs
Junyoung Park, Jungwook Choi, Mingu Lee
NLP Large Language Models Efficient ML
  • ValueDiff is a value-geometric eviction method that improves token retention in sink-suppressed LLMs.
  • The method scores tokens based on the dispersion of their value vectors, rather than relying on key-side information.
  • Empirical results show that ValueDiff achieves superior retention rates compared to prior eviction methods across multiple benchmarks.
  • The study highlights the importance of value geometry as a reliable eviction signal in modern LLM architectures.
Read more
G-NAC: Graph Neural Automata Clustering via Emergent Domain Formation
Keith Miller, Tristan Crawford
Graph Learning
  • Introduction of G-NAC, an unsupervised clustering framework leveraging recurrent graph-neural cellular dynamics.
  • Formulation of a domain-formation objective that encourages coherent representations without cluster-label supervision.
  • Development of a relational inference procedure based on graph-edge distance-rank stability.
  • Evaluation against multiple clustering baselines, demonstrating competitive performance and robustness.
Read more
$t_0$: A Time-Series Foundation Model for Forecasting with Context
Lucas Meyer, Claudio Sole, Huikan Xiang, Nicolas Li, Lucas Franceschino, Arnau Quera-Bofarull, Maarten P. Scholl, Joachim Fainberg, Geoffrey Négiar
Time Series
  • Introduction of $t_0$, a family of time-series foundation models for multivariate forecasting.
  • Models $t_0$-alpha and $t_0$-beta utilize transformer architecture with alternating attention mechanisms.
  • Significant performance improvements observed when using known-future covariates.
  • Models demonstrate strong performance on public benchmarks, ranking third in multiple evaluations.
Read more
NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware
Leonidas Delimpasis, Panagiota Moraiti, Antonis Porichis, Panos Chatzakos, Michail Karamousadakis
Multimodal Audio & Speech Robotics
  • Development of a neuromorphic AVSR system for noisy industrial environments.
  • Separation of spatial and temporal encoding into distinct modules for efficiency.
  • Significant reduction in word error rates compared to audio-only systems.
  • Energy-efficient operation on neuromorphic hardware with substantial power savings.
Read more
GraphToolbox: A Configurable Python Framework for Graph Neural Network Forecasting
Eloi Campagne, Yvenn Amara-Ouali, Yannig Goude, Argyris Kalogeratos
Graph Learning Time Series Interpretability
  • GraphToolbox unifies various stages of GNN forecasting into a single framework.
  • It supports data-driven graph construction and empirical selection methods.
  • The framework includes an adapter for 51 convolution operators and recurrent cells.
  • Online expert aggregation improves forecasting accuracy significantly.
Read more
Prescriptive SVD-Inspired Attention via Spectral Energy Retention
Vasileios Arampatzakis, Vasileios Sevetlidis, George Pavlidis
Computer Vision Interpretability Efficient ML
  • Introduces a diagnosis–intervention–verification framework for SVD-Inspired Attention (SVDA).
  • Demonstrates that spectral energy retention can effectively prune low-energy attention directions.
  • Empirical results show a reduction in parameters and computational costs without significant accuracy loss.
  • Establishes a structured approach to modify attention mechanisms based on spectral diagnostics.
Read more
CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning
Minkyoung Kim, Daeun Ji, Yohan Lee, Beomsoo Kim, Beakcheol Jang
Time Series Large Language Models
  • CTRL separates semantic reasoning from numerical prediction, enhancing stability in time series forecasting.
  • Specialized LLM agents analyze prediction errors through decomposed components, generating effective correction policies.
  • The framework allows for test-time adaptation without relying on ground truth, making it efficient in dynamic environments.
  • CTRL shows significant performance gains in non-stationary settings while remaining competitive in stationary scenarios.
Read more
COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules
Sushovan Majhi, Atish Mitra, Žiga Virk, Pramita Bagchi
Theory Optimization Efficient ML
  • Introduces COMPLEX, a closed-form embedding for multiparameter persistence modules.
  • Establishes the first two-sided distortion bound for multiparameter feature maps.
  • Achieves state-of-the-art performance on Orbit benchmarks without training.
  • Demonstrates that separated modules remain distinct in the embedding space.
Read more
SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras
Large Language Models Theory Generative Models
  • Introduces BENCHPROOFER, a pipeline for formal verification of LLM-generated code.
  • SWE-PROOF benchmark includes 500 real-world coding tasks with formal correctness proofs.
  • Demonstrates that many test-passing patches are flawed, emphasizing the limitations of traditional testing.
  • Finds that the quality of specifications significantly impacts verification success rates.
Read more
OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems
Kairui Yang, Xunkai Li, Kaixiang Zhang, Minghao An, Zekai Chen, Yuxuan Ba, Rong-Hua Li
NLP Large Language Models Graph Learning
  • OpenMAS-GCom provides a controlled environment for diagnosing performance in G-MAS by isolating specific components.
  • The benchmark evaluates 17 configurations across 29 datasets, introducing 400 complex tasks for comprehensive assessment.
  • Experiments reveal significant performance degradation when specialist agents are removed compared to critic agents.
  • Different configurations achieve varying levels of accuracy, highlighting the importance of communication structures and role assignments.
Read more
OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios
Yewen Li, Peng Jiang, Yitian Li, Pengfei Lv, Xialong Liu, Peng Jiang, Qingpeng Cai
Reinforcement Learning Generative Models Optimization
  • OneBid unifies multiple oCPX advertising scenarios into a single auto-bidding model.
  • The model employs a Mixture-of-Experts architecture to balance shared and scenario-specific knowledge.
  • Introduces a novel optimization method (CROP) for safe offline policy improvement.
  • Achieves significant performance gains in conversion scenarios, with an overall +2.2% improvement in ADVV.
Read more
PROSE: A Theory of Optimal Stopping with Perishable Evidence for Peer Selection in Intermittently Connected Decentralised Learning
Christos Anagnostopoulos
Federated Learning Theory Optimization
  • Develops a theory of optimal stopping for peer selection in decentralized federated learning.
  • Introduces a reservation-value threshold rule for optimal decision-making under perishable evidence.
  • Establishes confidence bounds and a maximin certification rule for peer selection.
  • Demonstrates the impact of mobility on the value of information and stopping decisions.
Read more
Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions
Zeyu Dong, Jiahui Zhong
Theory Optimization
  • High aggregate accuracy in single-cell models often masks poor performance on rare cell types.
  • Systematic benchmarking reveals that loss function optimization alone is insufficient for certain rare classes.
  • Absolute training-set size predicts the effectiveness of reweighting strategies, not relative frequency.
  • Class-balanced loss and LDAM are the most effective long-tail loss functions across various settings.
Read more
ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment
Yanxiao Liu, Sicheng Wan, Deniz Gündüz
Large Language Models Efficient ML Theory
  • Introduction of ExpBoN, a soft Best-of-n method using exponential noise for improved LLM alignment.
  • Theoretical guarantees of exponentially fast convergence and regret behavior for ExpBoN.
  • Integration of ExpBoN into the GSI framework, leading to significant computational savings.
  • Empirical results demonstrating substantial reductions in computational costs while preserving accuracy.
Read more
Dissecting Hierarchical Reasoning Models: A Mechanistic Study
Leo Raphael Rodrigues, Jian Kang
Theory Interpretability
  • HRM outperforms one-pass Transformer baselines in reasoning tasks.
  • The contributions of high- and low-level states in HRM vary by task and inference stage.
  • Linearly decodable features do not necessarily indicate causal relevance in HRM's reasoning process.
  • SAE ablations produce larger behavioral changes than linear probe direction ablations.
Read more
Token Utility Is Selection-Conditioned: Coupled Selection of Prompt Context and Response Supervision for Efficient Instruction Tuning
Can Wu, Xinrui Chen, Ou Wu, Yi Du
NLP Large Language Models Efficient ML
  • Introduces the concept of selection-conditioned token utility, emphasizing the interdependence of prompt and response selections.
  • Develops BRIDGE, a method that coordinates prompt and response selections through a shared validation-directed interaction surrogate.
  • Demonstrates that BRIDGE outperforms existing independent selection methods across multiple model families and tasks.
  • Finds that the performance advantage of BRIDGE increases with stronger compression in mathematical reasoning tasks.
Read more
Multi-Domain Clustering via Measure Quantization
Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante
Optimization Theory Multimodal
  • Introduction of a general framework for multi-domain clustering via measure quantization.
  • Utilization of probability metrics like Sinkhorn divergence and Maximum Mean Discrepancy for clustering.
  • Scalable mini-batch optimization strategy for efficient clustering.
  • Demonstrated superior performance on multiple benchmarks compared to classical methods.
Read more
From Regional to Global: Transfer Learning for Atmospheric Transport Emulators
Jeff Clark, Elena Fillola, Nawid Keshtmand, Raul Santos-Rodriguez, Matthew Rigby
Graph Learning Efficient ML Theory
  • Introduction of GATES, a machine learning emulator for atmospheric transport, which operates 1,000 times faster than traditional LPDMs.
  • Evaluation of model performance across four distinct global regions to assess spatial transferability and generalization capabilities.
  • Use of leave-one-region-out experiments to investigate the effectiveness of transfer learning in atmospheric transport modeling.
  • Characterization of regional differences in atmospheric transport to inform model training and improve global applicability.
Read more
RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks
Kai Ouyang, Dongyang Hou, Liangtian Liu, Zeyuan Wang, Ziyu Li, Chengfu Liu, Zichao Tang, Xuezhi Cui, Shengwu Ouyang, Wentao Yang, Hanwen Yu, Haifeng Li
Reinforcement Learning Large Language Models NLP
  • Introduction of RS-Claw-Evolution framework for lightweight RS agents.
  • Three-stage evolution process: interaction, experience, and decision evolution.
  • Implementation of programming-based interaction to manage states and reduce redundancy.
  • Feedback-driven fine-tuning strategy to enhance learning from errors.
Read more
Task-Aware Hybrid QUBO Optimization for Structured Neural Network Pruning
Osama Orabi, Artur Zagitov, Hadi Salloum, Viktor A. Lobachev, Yaroslav Kholodov
Optimization Efficient ML Computer Vision
  • Introduces a hybrid global-local framework for task-aware mixed-precision quantization.
  • Utilizes a Task-Aware QUBO formulation that incorporates quantization error and layer sensitivity.
  • Implements graph-aware structural constraints to enhance compatibility in activation precisions.
  • Demonstrates improved performance on a compact image-denoising model compared to existing methods.
Read more
Trading Depth for Time in Recurrent Transformers
Zeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen
NLP Large Language Models Efficient ML
  • Latent Recurrent Transformers (LRTs) introduce a thought token for hidden state refinement between vocabulary tokens.
  • The thought token model achieves comparable performance to deeper models while using approximately 48% fewer parameters.
  • Temporal computation through thought tokens can recover a significant portion of the benefits associated with increased physical depth.
  • The study provides insights into the depth-time trade-off in recurrent Transformer architectures.
Read more
GLR-MM: Graph-Based Global-Local Reconstruction for Robust Multimodal Chest X-ray and EHR Representation Learning under Missing Modalities
Surbhi Sharma, Nikhil Manali, Devesh Maheshwari
Multimodal Graph Learning Computer Vision
  • GLR-MM effectively reconstructs missing modalities using both local patient data and global cohort information.
  • The framework employs a graph-based approach to enhance the prediction of ICU mortality under varying levels of missing data.
  • GLR-MM shows superior performance compared to existing models, particularly in scenarios with high missingness (50%).
  • The model integrates adaptive fusion techniques to optimize the use of available data for improved predictions.
Read more
iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs
Ahmed Khaled Khamis, Xiaotong Ji, Hassan Jaber, Rasul Tutunov, Matthieu Zimmer, Jun Wang, Haitham Bou-Ammar
Large Language Models NLP Theory
  • iSDFT allows controlled information transfer from teacher to student models, improving continual learning.
  • The method introduces a new target-selection principle based on information constraints, enhancing local adaptation.
  • iSDFT consistently outperforms vanilla SDFT across multiple tasks and models, demonstrating its robustness.
  • The approach maintains a high retention rate of original capabilities while improving task-specific performance.
Read more
PAGE: Partition-Aware Gated KV-Cache Eviction
Pankaj Kumar, Subhankar Mishra
NLP Large Language Models Efficient ML
  • PAGE reframes KV-cache eviction as an input-dependent admission decision.
  • It identifies two classes of inputs: capacity-bound (harmful eviction) and dilution-prone (beneficial eviction).
  • The method utilizes a label-free scalar from prefill attention to predict eviction safety.
  • PAGE significantly reduces the harm rate of eviction from 0.75 to 0.026, a 29× reduction.
Read more
Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars
Daniel Henel, Frederik Werner, Alexander Langmann, Johannes Betz
NLP Audio & Speech Robotics
  • Jarvis is an offline voice assistant specifically designed for autonomous racing applications.
  • The framework integrates speech recognition and command classification, achieving low latency and high accuracy.
  • Experimental results demonstrate superior performance compared to larger online-hosted models.
  • The authors provide an open-source implementation to support further research.
Read more
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
Julien Siems, Riccardo Grazzi, Korbinian Pöppel, Jaisidh Singh, Arber Zela, Timur Carstensen, Jenia Jitsev, Frank Hutter, Volkan Cevher, Antonio Orvieto, Aaron Klein
NLP Large Language Models Efficient ML
  • CKDA enhances KDA's expressivity by enabling 2D rotations through a combination of delta-rule transformations and channel-wise reflections.
  • The model maintains non-expansiveness and computational efficiency while achieving the expressivity of DeltaProduct2.
  • CKDA can track finite groups isomorphic to subgroups of SO(3) with fewer layers than existing models.
  • Empirical results show CKDA outperforms Transformers and other linear RNNs in language modeling tasks.
Read more
AirGC-CD: Gaussian-Circulant Precoding for Exactly Debiasable PAPR Reduction in Over-the-Air Federated Learning
Jonggyu Jang, Hyeonsu Lyu, Hyun Jong Yang
Federated Learning Optimization Efficient ML
  • AirGC-CD effectively reduces PAPR in over-the-air federated learning without introducing bias.
  • The scheme utilizes Gaussian-circulant precoding to ensure unbiased aggregation of model updates.
  • It compresses data transmission efficiently, reducing the number of channel uses required.
  • The convergence analysis shows a favorable rate without a bias floor, enhancing learning performance.
Read more
Scaling Discovery through Test-Time Communication
Jongho Park, Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy, Dimitris Papailiopoulos
Optimization Large Language Models Theory
  • Test-time communication among agents can lead to significant performance improvements over independent attempts.
  • A team of k communicating agents can match the success rate of 4k independent agents, with benefits compounding as the team size increases.
  • Communicating agents achieved new state-of-the-art results in polyomino packing and MNIST classifier compression tasks.
  • The effectiveness of communication is contingent on having sufficient compute resources and clear feedback mechanisms.
Read more
Neural Cellular Automata Learn General Features in their Hidden Channels
Etienne Guichard, Stefano Nichele
Efficient ML Theory Computer Vision
  • NCAs provide a parameter-efficient alternative to traditional deep learning models, reducing the risk of overfitting.
  • The paper introduces a novel transfer-learning mechanism using hidden states instead of weights, enhancing few-shot learning capabilities.
  • NCAs outperform recurrent and feed-forward architectures on MNIST benchmarks with fewer than 10,000 parameters.
  • Hidden channels in NCAs capture general topological features, allowing for effective transfer of knowledge across classes.
Read more
Physics-residual machine learning predicts oxygen-evolution catalyst activity beyond the training range from sparse polarization measurements
Yong-Woon Kim, Jihyeok Lee, Sungtae Park, Sooseok Choi, Yung-Cheol Byun
Optimization Theory Efficient ML
  • Introduces a physics-residual machine learning approach for predicting catalyst activity.
  • Achieves significant error reduction compared to traditional data-driven models.
  • Demonstrates that a small number of labeled catalysts can effectively train the model.
  • Enables predictions beyond the training range, facilitating better candidate selection.
Read more
The Limits of Speculation: Bounding Speculative Decoding in Mixture-of-Experts
Aidar Amankulov, Denis Mamatin
NLP Large Language Models Optimization
  • Introduces the Sequence-Conditioned Offline Oracle methodology for estimating MoE verification costs.
  • Demonstrates diminishing returns when increasing the speculation budget in MoE models.
  • Establishes a linear relationship in Delta Space Analysis for decision-making in speculative decoding.
  • Provides a theoretical framework for understanding the limits of speculative decoding in MoE architectures.
Read more
SolarFlowRefiner: Refinement-Aware Flow Matching for Surface Solar Radiation Downscaling
Udbhav Srivastava, Antonita Racheal, Yiheng Chen, Runlong Yu, Xinyue Ye
Generative Models Optimization Time Series
  • Introduces SolarFlowRefiner, a refinement-aware framework for SSR downscaling.
  • Addresses the challenge of reconstructing high-resolution SSR fields from coarse ERA5 data.
  • Utilizes a conditional FlowMatch generator and a refiner trained on structured errors.
  • Demonstrates significant improvements over traditional standalone and post-hoc refinement methods.
Read more
A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare
Majid Lotfian Delouee, Sjors G. J. G. In 't Veld, Martijn C. Schut
Theory Interpretability Efficient ML
  • Introduction of OpTFM, a framework for evaluating foundation models on tabular data in healthcare.
  • Evaluation across six dimensions relevant to clinical applications, allowing for nuanced model comparisons.
  • Demonstration of varying model rankings based on specific healthcare use cases.
  • Provision of a taxonomy of 45 foundation models categorized by architecture.
Read more