AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

70 Papers today
8h Update frequency
7 Days of history
The Role of Feed-Forward Layers in Transformer Dynamics
Thomas Jacob Maranzatto, Semih Akkoc, Sennur Ulukus
Theory NLP Large Language Models
  • The feed-forward layer can steer tokens towards consensus in transformers, regardless of the attention matrices.
  • Theoretical results extend to multi-cluster convergence and multi-head attention scenarios.
  • Numerical experiments confirm the theoretical predictions and reveal richer dynamics in real-world LLMs.
  • The study connects transformer dynamics to control theory, providing a framework for future analysis.
Read more
RMB: Reward Model Boosting Mitigates Reward Hacking
Jiabin Fan, Dezhi Ye, Yongchang Hao, Lili Mou
Reinforcement Learning Large Language Models NLP
  • RMB addresses the reward hacking issue in RLHF by enhancing the robustness of reward signals.
  • The approach involves training multiple diverse reward models and aggregating their outputs using boosting techniques.
  • Extensive experiments show significant improvements in reward prediction accuracy and mitigation of reward hacking.
  • The use of a diversity-promoting regularizer helps ensure that reward models capture complementary aspects of the reward landscape.
Read more
Benchmarking graph-based models for in-silico toxicity prediction in drug discovery
Noel Suarez-Barro, Manuel Lama, Juan C. Vidal
Graph Learning
  • Introduction of a unified benchmarking framework for GDL-based toxicity prediction.
  • Systematic evaluation of over 20 models under consistent experimental conditions.
  • Fair and reproducible comparisons across multiple datasets and partitioning strategies.
  • Analysis of methodological trends and performance claims in existing literature.
Read more
Preferent Compression Bounds Are Tight
Dario Paccagnan, Marius Tirlea
Theory Optimization
  • The paper confirms that the state-of-the-art preferent compression bounds are tight.
  • An explicit construction using uniform distribution and order statistics is provided to demonstrate tightness.
  • A simpler proof of the upper bound is presented, making the results more accessible.
  • The findings have broad applications in risk certification across multiple domains, including machine learning and control systems.
Read more
Learning in the Transverse Subspace: A Minimal Representation for Divergence-Free Operator Learning
Yifei Sun
Theory
  • Introduces a minimal representation for divergence-free vector fields, reducing dimensionality from D to D-1.
  • The method ensures that the learned outputs are inherently divergence-free, eliminating the need for post-processing projections.
  • Utilizes Fourier extension for nonperiodic flows to maintain compatibility with divergence-free constraints.
  • Demonstrates improved performance in terms of error reduction and robustness in operator learning tasks.
Read more
GNA: Granular Neighbor Assembly for Retrieval-Augmented Multivariate Time-Series Forecasting
Vincent Uhse
Time Series
  • GNA combines whole-window and per-variate retrieval in a single multivariate forecasting model.
  • The model uses a learned gate to weigh the trust in retrieved futures against persistence forecasts.
  • GNA shows significant performance improvements across multiple datasets and benchmarks.
  • Retrieval is most effective when the lookback window is less informative, allowing the model to adaptively shift trust.
Read more
AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin, Yao Lai, Haoran Wu, Nicholas D. Lane, Robert D. Mullins, Ilia Shumailov, Yiren Zhao
Large Language Models
  • AgentPerfBench captures diverse agentic workloads, including coding and tool-using agents, through multi-turn interactions.
  • The benchmarking suite employs saturation-based measurements to reflect true hardware performance under maximum load.
  • Kernel-level profiling reveals performance bottlenecks and provides insights into memory bandwidth and capacity limitations.
  • The study demonstrates that existing benchmarks fail to accurately represent the demands of agentic LLM applications.
Read more
Identifying ODEs from Unstructured Data with Causal Representation Learning
Alessandro Trenta, Riccardo Massidda, Davide Bacciu, Sara Magliacane
Theory Time Series Computer Vision
  • Introduces SPEED-AE, a framework combining CRL and autoencoders for ODE discovery.
  • Demonstrates improved identifiability of variables from polynomial to monomial diffeomorphisms.
  • Achieves state-of-the-art performance in recovering ODEs from unstructured data.
  • Provides theoretical guarantees for the identification of true variables from high-dimensional observations.
Read more
ConRAG: Lightweight inference of multi-hop relations
Kilian Bรคnziger, Sonia Laguna, Markus Kreft, Robert Jakob, Kevin O'Sullivan, Lasse B. Strand, Julia E. Vogt
NLP Large Language Models Graph Learning
  • Introduces CONRAG, a graph-based framework for multi-hop relation inference.
  • Constructs a lightweight entity-document graph for efficient path retrieval.
  • Outperforms existing RAG baselines in bridge entity recovery and reasoning chain precision.
  • Reduces indexing token costs significantly, enhancing scalability.
Read more
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
Oliver Sieberling, Bharat Runwal, David Jin, Ryan Chin, Rameswar Panda, Yoon Kim
NLP Large Language Models Efficient ML
  • Triadic Linear Attention increases the memory state size of RNNs using a triadic outer product, enhancing recall for long-context tasks.
  • The method is parameter-efficient, allowing for significant state size increases with minimal additional projections.
  • Compatible with modern innovations in linear attention, such as data-dependent forgetting and the delta rule.
  • Demonstrated improvements in long-context language modeling and recall capabilities over traditional and alternative approaches.
Read more
Inducing Process Supervision from Outcome-Only Reinforcement Learning
Shengda Fan, Xin Cong, Zhong Zhang, Haotian Chen, Yankai Lin
Reinforcement Learning Large Language Models Theory
  • Introduction of TIPS, an outcome-only RL framework for training PRMs.
  • TIPS reinforces step-level verification through outcome prediction without direct supervision.
  • Achieves state-of-the-art performance on ProcessBench with significantly fewer labeled trajectories.
  • Provides a theoretical analysis supporting the relationship between outcome verification and step correctness.
Read more
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
Yuchen Cai, Ding Cao, Qixiang Yin, Xin Xu, Kai Yang, Siye Wu, Pengyuan Wang, Jiaxuan Wang, Weijie Liu, Saiyong Yang, Guangzhong Sun, Guiquan Liu, Junfeng Fang
NLP Large Language Models Reinforcement Learning
  • Identification of a low-dimensional effective manifold in activation space related to RL-induced performance gains.
  • Discovery of two geometric properties: Effective Manifold Capacity and Control Manifold Separation.
  • Introduction of Alpha-Stabler, a framework that stabilizes RL training and enhances performance.
  • Experiments demonstrate that a single input-invariant vector can recover a significant portion of RL gains.
Read more
Learning the Structure of Triangular Transport Maps
Morten Blรธrstad, Pekka Parviainen, Berent ร…nund Strรธmnes Lunde
Generative Models Graph Learning Optimization
  • Introduces Self-Structuring Transport Maps (SSTM) for joint learning of map structure and parameters.
  • Utilizes SoftSort for variable ordering and stochastic L0 gates for sparsity learning.
  • Demonstrates superior density estimation performance compared to traditional methods.
  • Achieves competitive results against autoregressive flows on large datasets.
Read more
Neural Succession: A Mesoscopic Theory of Invasion, Coexistence, and Stabilization in Continual Learning
Shoaib Ahmed Dipu, Md Salman Shamil, Sayeed Shafayet Chowdhury
Theory
  • Introduces Successional Learning Theory (SLT) as a mesoscopic framework for continual learning.
  • Demonstrates that pre-invasion compatibility predicts forgetting and coexistence outcomes effectively.
  • Establishes ecological analogies for understanding task transitions in continual learning.
  • Identifies key conditions for coexistence and provides a minimum habitat-modification bound.
Read more
Variational Augmented Invertible Koopman Autoencoder for probabilistic time series forecasting
Anthony Frion, Lucas Drumetz, Guillaume Tochon, Mauro Dalla Mura, Ali Can Bekar, Abdeldjalil Aรฏssa El Bey
Time Series
  • VAIKAE introduces a stochastic framework for time series forecasting, enhancing uncertainty quantification.
  • The architecture leverages normalizing flow models for likelihood computations in dynamical systems.
  • New strategies for uncertainty-aware latent data assimilation are proposed.
  • Experiments show VAIKAE's effectiveness on long-term forecasting benchmarks.
Read more
Depot-Closed Multi-Component Construction for Neural Vehicle Routing
Shinichiro Hamada, Hisashi Kashima
Optimization
  • Introduction of multi-component construction for flexible route assignment in vehicle routing.
  • Depot-closed interpretation allows for effective evaluation of route components and merges.
  • Neural policy trained on CVRP100 shows strong generalization to larger problem sizes.
  • Outperforms existing neural solvers in various benchmarks, including zero-shot evaluations.
Read more
BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions
Thanina Hamitouch, Khadidja Henni, Abdelkrim Arie, Amina Selma Haichour, Neila Mezghani, Lina Abou-Abbas
NLP Large Language Models Multimodal
  • BERT4DTI combines ChemBERTa and ProtBERT with mutual attention for DTI prediction.
  • The model addresses limitations of existing DTI models, such as data scarcity and computational cost.
  • BERT4DTI achieves state-of-the-art results on multiple benchmark datasets.
  • The use of partial fine-tuning significantly reduces the number of trainable parameters.
Read more
Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space
Tue M. Cao, Lisiane Pruinelli, My T. Thai
Large Language Models Interpretability
  • Introduction of ASPD for joint decomposition of activation and parameter spaces.
  • Grounding weight components in activation features enhances interpretability.
  • ASPD enables scalable and causally editable parameter decomposition in large models.
  • Demonstrated effectiveness on Qwen-3-8B, recovering known computational mechanisms.
Read more
Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
Haozhan Tang, Hao Kang, Han Cai, Song Han, Chenyan Xiong
NLP Large Language Models Efficient ML
  • Identification of stale-delta failure mechanism in FP8 attention training.
  • Introduction of Delta-Matching to restore softmax gradient invariance.
  • Demonstration of Delta-Matching's effectiveness across various model sizes and architectures.
  • Matching of BF16/FP32 mixed-precision training performance with native FP8 training.
Read more
Transversal Pooling Neural Networks
Emily J. King, Dustin G. Mixon, Michael Perlmutter, Lander Ver Hoef
Computer Vision Theory Efficient ML
  • Introduction of Transversal Pooling Neural Networks (TraPNets) for improved stability and sensitivity to transformations.
  • Establishment of equivariance to affine group actions and derivation of stability bounds for pooled coefficients.
  • Demonstration of TraPNets' effectiveness in low-data scenarios, particularly in tropical cyclone prediction.
  • TraPNets outperform traditional CNNs with data augmentation in synthetic experiments.
Read more
Interpretable intrinsic dimension estimation through componentwise calibration of distance and angle
Chih-Hsuan Huang, Chih-Wei Chen, Szu-Chi Chung
Interpretability
  • Introduces componentwise calibration for intrinsic dimension estimation, enhancing interpretability.
  • Derives a closed-form Kullback-Leibler divergence for improved robustness against noise.
  • Demonstrates significant reductions in mean percentage error in ID estimation across various datasets.
  • Highlights the impact of sample-amplitude heterogeneity on angular statistics.
Read more
RAEGL: Risk-Aware Evidence-Gated Learning for Selective Contextual Routing under Temporal Shift
Yifan Guo
Time Series
  • RAEGL introduces a deploy-or-exact-fallback approach for contextual forecasting under temporal shifts.
  • The framework separates model training, candidate selection, gate calibration, and evaluation into distinct phases.
  • A search-aware evidence gate ensures that contextual components are only deployed when they meet specific criteria.
  • Experiments show RAEGL's effectiveness in reducing forecasting errors and managing deployment risks.
Read more
Equivariant Neural Primal-Dual Assignment for Maximum Common Edge Subgraphs
Jiaqing Xie, Yanchao Li, Zhuo Yang, Yuxin Wang, Tianfan Fu, Yuqiang Li
Graph Learning Optimization Theory
  • ENPDA allows for a reusable matching policy that significantly reduces query time for MCES problems.
  • The method provides per-pair guarantees and optimality bounds, ensuring reliable performance across different graph pairs.
  • ENPDA outperforms traditional methods by 7.4-8.6 accuracy points on molecular benchmarks and shows strong transferability to other graph tasks.
  • The approach maintains a frozen network structure while adapting bids and prices, optimizing the matching process efficiently.
Read more
EvoMO-SR: Multiobjective LLM-based Evolution of Symbolic Expressions with substructure guidance
Cristina Rossetti, Anna V. Kononova, Thomas Bรคck, Fei Liu, Niki Van Stein
Large Language Models Optimization Interpretability
  • Introduction of EvoMO-SR, an LLM-driven framework for Symbolic Regression.
  • Implementation of a multi-objective survival selection to manage formula complexity and accuracy.
  • Use of a substructure guidance mechanism to enhance expression mutation.
  • Demonstrated superior performance on LSR-Synth and custom datasets compared to traditional SR methods.
Read more
When Trees Are Not Enough: Learning Mixed-Topology Feature Graphs with Adaptive Graph Sparse Autoencoders
Xiaozuo Shen, Yifei Cai, Tian Tan, Rui Ning, Chunsheng Xin, Hongyi Wu
NLP Large Language Models Graph Learning
  • AG-SAE allows for mixed-topology feature graphs, overcoming the limitations of single-parent tree structures.
  • The framework identifies necessary multi-parent relations while rejecting redundant or spurious alternatives.
  • AG-SAE employs a self-consistency cycle between dictionary and graph learning for continuous refinement.
  • Experimental results show improved relational reliability and semantic validity over existing methods.
Read more
NeuronSifter: Intervention Planning in CNS Microenvironments
Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie
Optimization Theory
  • NeuronSifter integrates structured intervention compilation with decision-directed evidence acquisition.
  • The framework improves intervention ordering accuracy from 0.760 to 0.880 in synthetic AD evaluations.
  • It retains joint uncertainty across candidate rollouts, enhancing decision quality.
  • NeuronSifter demonstrates superior performance compared to traditional Bayesian experimental design planners.
Read more
Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase
Zhang Yanhai
Theory Efficient ML Optimization
  • Introduces a continual learning framework that operates without an offline phase.
  • Utilizes biologically inspired mechanisms such as k-winner-take-all dynamics and refractory rotation.
  • Achieves competitive performance on benchmark datasets, surpassing traditional methods in certain conditions.
  • Demonstrates the potential for memory consolidation during active inference rather than requiring dedicated offline processing.
Read more
Replication Failure and Trivial Baselines in Road-Level Crash Prediction
Maurya Patel
Graph Learning
  • Only 4 out of 11 design decisions from a previous model replicate in a second borough.
  • Multi-seed evaluation reveals significant variability in model performance.
  • The proposed GNN model is statistically indistinguishable from a trivial baseline based on past crash counts.
  • Published results may overstate the advantages of GNNs due to short lookback horizons.
Read more
Why Backdooring Neural Networks is so Easy?
Issam Seddik, Mohamed El Amine Seddik
Theory
  • Feature learning in neural networks increases vulnerability to backdoor attacks.
  • The relationship between poison fraction and trigger strength is characterized by ฮฑ โˆ ฯ€โˆ’1/4 in feature-learning regimes.
  • Existing security audits may underestimate backdoor risks due to reliance on linear heuristics.
  • The study provides a theoretical framework that aligns with large-scale empirical findings.
Read more
Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution
Yubin Lyu, Fu Li, Jiawei Fei, Yang Zhao, Weixing Mei, Yinan Wu
Optimization
  • Mara Chain retains and refines rejected candidates instead of discarding them, leveraging past failures for future improvements.
  • The method limits the depth of refinement chains and employs Pareto-filtered Top-N selection to manage candidate pools effectively.
  • Mara Chain outperforms existing optimization methods by significant margins, achieving better results with fewer rollouts.
  • The approach is validated across multiple benchmarks, demonstrating its versatility across different types of AI artifacts.
Read more
Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS
Paweล‚ Lenartowicz, Hubert Plisiecki
Theory Efficient ML Interpretability
  • Introduction of two new tests for PLS inference: a corrected t-test and a permutation test.
  • Demonstration of significant power improvements over existing methods like CV-permutation-Q2.
  • Validation of methods on various datasets, including synthetic and real-world applications.
  • Establishment of a framework for per-component inference in PLS models.
Read more
DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation
Dwipam Katariya, Thomas Caputo, Akshat Shreemali, Juan Manuel Origgi, Nikita Seleznev, Pranab Mohanty, Kalanand Mishra, Nam Nguyen, James Montgomery
Efficient ML Time Series NLP
  • DP-Rec is the first to apply dynamic latent patching for sequential recommendation, achieving superior efficiency and accuracy.
  • Introduces Contrastive Entropy Surprise as a computationally efficient method for determining dynamic behavioral boundaries.
  • Incorporates temporal dynamics as first-class features to enhance the detection of informative segment boundaries.
  • Demonstrates significant improvements in performance under constrained computational budgets across multiple datasets.
Read more
Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks
Dongyub Jude Lee, Jungseob Lee, Chanjun Park, Hyeonseok Moon, Heuiseok Lim
Large Language Models NLP Theory
  • PriceCheck introduces a pricing mechanism for label-free checks to control selective risk in LLM answering.
  • The method achieves an average of 76.1% answer serving while keeping selective risk below 1.5%.
  • Price-based coverage predictions show a high correlation (0.97) with observed coverage across various schedules.
  • PriceCheck outperforms traditional methods, including reward models and correctness classifiers, in terms of serving more answers with fewer errors.
Read more
Selective Elicitation as a Commercial Influence Channel: A Reproducible Synthetic Shopping-Agent Stress Test
Jiapeng Li
Theory
  • Commercial incentives can influence the questions asked by AI shopping assistants, affecting user responses.
  • The study contrasts neutral, soft commercial, and targeted questioning policies in a synthetic setting.
  • Targeted questioning significantly increases the selection of sponsored products while reducing overall utility.
  • The findings highlight the need for careful consideration of question policies in AI recommendation systems.
Read more
Shared Autoregressive Context Can Distort Relationships in Synthetic Data
Thomas S. Robinson
Large Language Models Generative Models Theory
  • Shared autoregressive context can significantly distort relationships in synthetic data generated by LLMs.
  • Generating multiple respondents in one completion increases mean absolute error in correlations compared to generating them separately.
  • Answer history acts as a causal channel influencing the relationships among generated responses.
  • Hiding preceding answers can reduce correlation errors but may worsen marginal accuracy.
Read more
Arithmetic Simplicity in Stochastic Gradient Methods
Bin Fu, Pengfei Gu, Jose Nunez, Fabian Vazquez
Optimization Efficient ML Theory
  • Introduces the concept of arithmetic simplicity in gradient descent methods.
  • Transforms AdaGrad, Adam, and AdamW into arithmetically simple versions suitable for hardware implementation.
  • Demonstrates faster convergence for the static version of Adam in experimental results.
  • Highlights the benefits of reduced hardware complexity and energy consumption.
Read more
Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories
Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu, Zhiqiang Zhang, Xiaolin Huang, Jun Zhou
Large Language Models NLP Optimization
  • Mid-training performance in LLMs saturates with increased compute, necessitating alternative allocation strategies.
  • Trajectory Soup enables the distribution of compute across multiple independent trajectories, enhancing model performance.
  • Inter-trajectory averaging significantly reduces residual error compared to traditional intra-trajectory methods.
  • The method shows improved performance across different model scales, learning rates, and token budgets.
Read more
Behavioral Capacity Certificates for Quantized Language Models
Arian Eamaz, Mojtaba Soltanalian
NLP Large Language Models Efficient ML
  • Introduces Behavioral Capacity Certificates (BCC) for quantifying model behavior in quantized language models.
  • Proposes a three-step workflow for model deployment that includes screening, certification, and bounding of population loss.
  • Demonstrates that BCC can lower complexity penalties while preserving model performance.
  • Finds that higher precision for keys than values in cache memory improves model predictions.
Read more
Strict-Saddle Landscapes and Multi-Rank Geometry in Low-Tubal-Rank Tensor Sensing
Eugene Agyei-Kodie, Longxiu Huang, Shuang Li, Xiao Liang
Optimization Theory
  • Establishment of a strict-saddle landscape for low-tubal-rank tensor sensing with no spurious local minima.
  • Local geometry is determined by Fourier-slice ranks rather than solely by tubal rank.
  • Uniform ranks yield quadratic growth, while nonuniform ranks lead to quartically flat directions.
  • Numerical experiments confirm the global optimization behavior and highlight differences in local geometries.
Read more
When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models
Yucong Cao, Chenqi Li, Tingting Zhu
Interpretability Time Series
  • The initial interpretation of SAE latents as representing alpha activity is misleading and largely due to input distortion effects.
  • A proposed validation ladder helps assess the interpretability of SAE features, emphasizing the need for rigorous testing.
  • The study reveals that latents selected for their response to alpha removal do not correlate with clean EEG alpha power.
  • The findings challenge existing assumptions about the semantic identifiability of features in EEG foundation models.
Read more
Multi-Agent Flow Matching with Decoupled Generative Guidance
Ruoyu Lin, Magnus Egerstedt, Fabio Pasqualetti
Generative Models Robotics Theory
  • Introduces DeGG-Flow for multi-agent flow matching with decoupled generative guidance.
  • Establishes formal guarantees for both shared and private requirements in multi-agent systems.
  • Provides feasibility and finite-horizon convergence guarantees for generated outputs.
  • Derives a Wasserstein bound to characterize distributional deviation under guidance.
Read more
Predictive Dual Smoothing for Column Generation
Senne Berden, Noah Schutte, Andrea Lodi, Tias Guns
Optimization
  • Introduction of predictive dual smoothing that uses future dual predictions for improved pricing.
  • The method combines current dual solutions with learned predictions to enhance column generation.
  • Extensive evaluation shows substantial efficiency gains in solving large-scale linear programs.
Read more
HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation
Yansong Liu, Rui Liu, Yuan Zuo, Hongwei Zhao, Da Fu, Fuwei Zhang, Fuzhen Zhuang, Yong Chen, Zhe Li
Graph Learning Time Series Theory
  • Identifies limitations of fixed-prefix protocols in extrapolative TKGR.
  • Reformulates TKGR as continual learning over streaming snapshots.
  • Proposes HiTS-CL, which combines continual fine-tuning, adaptive distillation, and selective memory.
  • Demonstrates consistent improvements in extrapolation accuracy across multiple datasets.
Read more
Correct then Forecast: Observer State-Space Models for Time Series Forecasting
Alexis-Raja Brachet, Guillaume Clavier--Frรฉmond, Abdelhakim Ziani, Pierre-Yves Richard, Cรฉline Hudelot
Time Series
  • OSSMs separate latent-state propagation from measurement assimilation, improving time series forecasting.
  • The framework allows for a unified interpretation of existing state-space models and reveals modeling inconsistencies.
  • OSSMs achieve substantial performance improvements while keeping the same parameter count as traditional models.
  • The proposed method emphasizes the role of observations in correcting estimated latent states rather than controlling dynamics.
Read more
Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training
Emre Can Acikgoz, Yang Li, Zeyu Leo Liu, Srijan Bansal, Dilek Hakkani-Tรผr, Shafiq Joty, Semih Yavuz
Large Language Models Reinforcement Learning NLP
  • The effectiveness of OPD is dependent on student-teacher compatibility rather than teacher scale alone.
  • A brief SFT warm-up improves OPD performance, while RLVR-prepared students may regress under distillation.
  • Adapting the teacher with RLVR enhances downstream OPD accuracy.
  • OPD offers a better initialization for subsequent RLVR compared to SFT at similar starting accuracies.
Read more
Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training
Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Grรฉgoire Danoy, Pascal Bouvry
Reinforcement Learning Large Language Models Optimization
  • Critic instability is an optimization artifact, not an inherent flaw.
  • RFPO utilizes a single frozen critic for multiple roles, enhancing efficiency.
  • Binarizing the critic's score mitigates length bias in policy optimization.
  • RFPO achieves performance parity with supervised PPO while reducing resource requirements.
Read more
Routing in Gradient Space: Balanced Usage Is Not Expert Specialization
Yuchen Li, Mingyu Du, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong
NLP Large Language Models Optimization
  • Introduction of Gradient-Aligned Routing (GAR) for optimizing expert assignments based on gradient coherence.
  • GAR outperforms traditional routing methods in multi-task text classification, achieving higher accuracy and better expert load balance.
  • The study distinguishes between expert load balance and gradient-based routing organization, emphasizing their separate impacts on model performance.
  • The methodology leverages a load-normalized partitioning criterion that enhances the coherence of gradient contributions to experts.
Read more
In-Context Learning Amplifies a Latent Symbolic Circuit
Melissa Wessel
NLP Large Language Models Interpretability
  • The symbolic reasoning circuit is identifiable and functional before achieving high accuracy.
  • Causal contributions from the circuit increase significantly with the number of in-context examples.
  • Patching higher shot activations into lower shot prompts can enhance model accuracy substantially.
  • Function vectors can effectively substitute for the induction stage in the reasoning process.
Read more
Seeing Time: Visual-Temporal Representation Learning for Interpretable Time Series Clustering
Zheng Zhu, Zexi Tan, Yuming Deng, Yiqun Zhang
Time Series Interpretability Multimodal
  • WAVE integrates temporal and visual representations for improved interpretability in time series clustering.
  • The method employs cross-modal contrastive learning to align different representation modalities.
  • WAVE achieves the highest macro-averaged clustering performance across multiple datasets.
  • The approach allows for waveform-level traceability, enhancing the validation of clustering results.
Read more
From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning
Yupeng Chang, Wenxuan Zhang, Yuan Wu
NLP Large Language Models Optimization
  • Checkpoint selection is crucial in supervised fine-tuning, impacting model performance significantly.
  • Increasing the validation budget improves selection accuracy, with notable gains observed in generated-accuracy and checkpoint-agreement methods.
  • Alternative selection rules can outperform traditional validation-loss selection but do not consistently outperform the final checkpoint.
  • The study highlights the need for distinct evidence to support claims regarding checkpoint selection performance.
Read more
TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
Kasra Arabi, Nir Weinberger, Micah Goldblum, Niv Cohen
NLP Large Language Models Generative Models
  • TANGO introduces a coloring-based watermarking method for masked-diffusion language models.
  • The watermark is embedded in pairs of tokens, reducing the risk of frequency-based attacks.
  • Detection does not rely on the order of unmasking and requires only the text and a secret key.
  • TANGO shows improved detection rates compared to traditional methods like the red-green list.
Read more
Adapting Linear-Time Architectures for Tabular In-Context Learning
David Schnurr, Felix Sarnthein, Thomas Hofmann, Imanol Schlag
Efficient ML
  • Identified optimal training strategies for causal and non-causal models in tabular ICL.
  • DeltaNet outperforms non-causal linear attention but struggles with longer contexts.
  • Introduced a decay schedule to stabilize recurrent state and improve generalization.
  • Proposed a final-state reading mechanism to enhance performance on benchmark datasets.
Read more
When the Merge Coefficient Stops Mattering: Proximity Regularized Merging for Continual LoRA Adaptation
Yixuan Liu, Yuhao Sun, Sen Song, Jin Li
Efficient ML Theory
  • Proximity Regularized Merging (PRM) enhances rehearsal-free continual learning by adding a proximal penalty during task-vector training.
  • The effectiveness of merging task vectors is influenced by their training conditions, not just the merge coefficient.
  • PRM consistently improves performance across multiple write-in rules and architectures, achieving the best reported mean AAA.
  • Mechanistic analyses show that PRM reduces Fisher-weighted interference and broadens the coefficient plateau.
Read more
Neural Constitutive Learning for Generalized Reaction-Diffusion Systems
Shang-Ke Chen, Yu-Peng Wang, Shih-Hsuan Hung, Wei-Fang Sun, Chao-Shun Zhan, Simon See, Min-Jhe Lu
Theory
  • Introduction of the NCL-MCT Solver for generalized reaction-diffusion systems.
  • Separation of PDE-specific constitutive responses from shared temporal evolution.
  • Support for trajectory-free constitutive learning and reuse across different conditions.
  • Demonstrated effectiveness across seven systems with competitive error rates.
Read more
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang
NLP Large Language Models Reinforcement Learning
  • Introduction of Least-Square Policy Distillation (LSPD) for improved sample efficiency in language model reasoning.
  • LSPD connects reverse-KL objectives in OPD with KL-regularized policy optimization, enabling off-policy data reuse.
  • Empirical results demonstrate significant performance improvements over existing distillation baselines.
  • LSPD maintains policy diversity and efficiency through exploration and historical data reuse.
Read more
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
Zhiwei Zhang, Huayu Deng, Fei Zhao, Jiayan Fu, Bin Liang, Kam-Fai Wong, Mu Chuan
Reinforcement Learning Large Language Models NLP
  • ROSS preserves historical self-generated rollouts as a reusable training resource.
  • The method selectively supervises informative segments of rollouts while maintaining full trajectory context.
  • ROSS shows consistent performance improvements across multiple domains, including math and coding tasks.
  • The approach highlights the importance of compatibility and complementarity in historical experiences for effective training.
Read more
Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos
Simeon Allmendinger, Domenique Zipperling, Burhanettin Bahadir Kibar, Niklas Kรผhl
Federated Learning Generative Models
  • CollaFuse enables collaborative synthetic data generation without direct data sharing.
  • The method addresses the scarcity of informative minority-class observations in fraud detection.
  • Synthetic data generated through CollaFuse improves fraud detection performance across multiple classifiers.
  • The approach reduces computational burdens on individual organizations compared to traditional federated learning.
Read more
Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
Muhang Tian, Sherry Yang
Reinforcement Learning Efficient ML Optimization
  • Introduces Reward-rate Policy Gradient (RPG) to optimize reward per unit of time in RL tasks.
  • RPG uses off-policy samples to estimate reward rates, avoiding costly on-policy rollouts.
  • Theoretical analysis shows RPG approximates optimal reward rates effectively.
  • Empirical results demonstrate significant performance improvements over traditional RL methods.
Read more
On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models
Shuoyuan Wang, Beier Luo, Hao Zeng, Chengyao Yu, Songxin Zhang, Zejian Xie, Bingyi Jing, Hongxin Wei
Reinforcement Learning Large Language Models Optimization
  • LRMs often exhibit a concentrated confidence prior that limits exploration in on-policy RL.
  • The proposed CalibSFT method reshapes the confidence prior to enable better calibration.
  • CalibSFT combines success rates with correctness to create balanced confidence targets.
  • Extensive evaluations show that CalibSFT improves calibration and discrimination without sacrificing accuracy.
Read more
MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models
Zhi Wen Soi, Giulio Segalini, Jian-Jia Chen, Lydia Chen
Multimodal Audio & Speech Large Language Models
  • Introduction of MISHAP-Bench, a benchmark specifically for evaluating hallucinations in LALMs.
  • Definition of two categories of hallucination: context and knowledge.
  • Development of a groundedness judge for evaluating open-ended responses.
  • Significant hallucination rates observed in state-of-the-art LALMs, highlighting the need for better mitigation strategies.
Read more
What You Observe Determines How You Identify Causal Effects: Evaluating Causal Models across Observational Views
Heejin Jung, Gyeongdeok Seo, Hoyoon Byun, Joseph Lee, Kyungwoo Song
Theory
  • Introduction of CausalIDView, a benchmark for evaluating causal estimators across different observational views.
  • No CFM consistently outperforms others; model performance varies significantly depending on the observational view.
  • Modular estimators combining predictive models with explicit identification procedures can achieve competitive results.
  • CFMs exhibit model-specific failures under structural changes, impacting the stability of causal effect estimates.
Read more
JudgeCast: Time Series Forecasting with Experience-Informed Covariate Judgements
Donguk Kwon, Wooseok Jeong, Dongha Lee
Time Series
  • JudgeCast introduces an experience-based framework for time series forecasting that utilizes covariate judgments.
  • The framework separates the assessment of covariate effects from the numerical adjustment process.
  • Residual-guided experience construction allows for the reconstruction of judgments based on forecast errors.
  • JudgeCast outperforms traditional forecasting methods and strong baselines in real-world datasets.
Read more
Rethinking Soft Tokens for Parallel Decoding in Diffusion Language Models
Kodai Kawamura, Kenji Kawaguchi, Anji Liu
NLP Generative Models Large Language Models
  • Soft tokens can mitigate inconsistencies in token generation during parallel decoding in DLMs.
  • A geometry-aware construction of soft tokens is proposed for use with frozen pretrained models.
  • Soft-token feedback enhances coherence in token sequences beyond just preserving uncertainty.
  • The method shows superior performance compared to existing parallel decoding strategies.
Read more
TokenCast: Forecasting Token Consumption During LLM Agent Execution
Chaoqian Ouyang, Ling Yue, Libin Zheng, Huanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
Large Language Models NLP Efficient ML
  • TokenCast introduces a segment-cost factorization for accurate token consumption forecasting.
  • The method updates predictions dynamically during execution without additional LLM calls.
  • TokenCast achieves a 14.5% reduction in mean absolute error compared to existing methods.
  • In budget-control scenarios, TokenCast uses 21.3% fewer tokens than fixed-budget policies.
Read more
Sharp Convergence and Sampling Trade-offs for Riemannian Diffusion under Nonnegative Ricci Curvature
Yuhao Liu, Longbo Huang
Generative Models Theory Efficient ML
  • Developed a general framework for Riemannian diffusion models that separates score discretization and Brownian-motion simulation errors.
  • Achieved a convergence rate of O(d/ฮตยฒ) for score evaluations under nonnegative Ricci curvature, matching Euclidean models.
  • Showed that O(dโดT/ฮตยฒ) geodesic random-walk steps suffice for accurate Brownian motion approximation.
  • Provided a strategy for performing multiple Brownian-motion simulation steps per score evaluation to optimize computational resources.
Read more
LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
Hao-Yuan He, Peng-Fei Liu, Si Shen, Ming Li
NLP Large Language Models Efficient ML
  • Introduces the concept of fixed-cost drafting, which allows drafter costs to remain constant regardless of context length.
  • LONGSPARK employs a block-diffusion approach to extract fixed-size, multiscale context views from the target model's verification pass.
  • Achieves significant improvements in end-to-end throughput and reduces drafter memory overhead by several orders of magnitude.
  • Demonstrates superior performance in long-context tasks, achieving the lowest time-per-output-token.
Read more
AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models
Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi
Multimodal Large Language Models Optimization
  • AdaKerNet operates on frozen MLLM representations, avoiding the need for parameter fine-tuning.
  • The architecture combines learnable multimodal features, a reference kernel, and a nonlinear predictor for task adaptation.
  • Empirical results show significant performance improvements over traditional decoding methods in limited supervision scenarios.
  • The framework allows for joint optimization of kernel representation and prediction, enhancing adaptability to downstream tasks.
Read more
Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space
Yuchen Li, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong
Optimization NLP Computer Vision
  • BOM replaces parameter-space momentum history with a compact task-space EMA, significantly reducing memory usage.
  • The method preserves the current supervised gradient while reprojecting historical information through the current model.
  • BOM achieves substantial improvements in validation performance across multiple tasks and models.
  • The approach is compatible with existing adaptive optimizers, enhancing their efficiency without compromising performance.
Read more
Interactive-Policy Distillation with Bidirectional Propose-and-Verify
Shutong Wu, Xiwen Chen, Brendan Rappazzo, Daiheng Zhang, Anderson Schneider, Yuriy Nevmyvaka, Jiawei Zhang
Large Language Models NLP Efficient ML
  • Introduction of Interactive-Policy Distillation (IPD) to improve knowledge distillation.
  • Bidirectional propose-and-verify mechanism allows for adaptive teacher intervention.
  • IPD outperforms traditional on-policy distillation methods in accuracy and data efficiency.
  • Fused inference engine designed to optimize dual-model rollouts.
Read more
Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs
M. Saeid HaghighiFard, Sinem Coleri
Federated Learning
  • Introduces EN-HMTFL, a framework for multi-task learning in VANETs that allows vehicles to share a common encoder while retaining local decoders.
  • Addresses the scalability and communication efficiency issues of traditional federated learning in dynamic vehicular environments.
  • Demonstrates up to 24% improvement in accuracy and a reduction of up to 69 communication rounds for convergence compared to benchmarks.
  • Utilizes a cluster-based hierarchical federated learning architecture to facilitate local model exchanges and reduce infrastructure load.
Read more