AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

66 Papers today
8h Update frequency
7 Days of history
Fisher Widths: Local Learning Geometry and Anisotropic Recovery
Vu Khac Ky
Theory Optimization
  • Introduction of primal and inverse Fisher widths as complementary functionals in statistical learning and recovery.
  • Demonstration of the relationship between local parameter fluctuations and recovery in the context of Fisher geometry.
  • Establishment of a two-sided estimate for statistical dimension and support-sensitive recovery estimates.
  • Identification of the influence of Fisher anisotropy on the complexity of learning and recovery tasks.
Read more
GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries
Miguel P. Bento, João Seabra
Large Language Models Efficient ML Optimization
  • GaugeQuant introduces a novel in-training quantization method that mitigates outlier effects without requiring calibration data.
  • The method modifies the loss function to include a LogSumExp term, effectively minimizing activation outliers during training.
  • It employs a stop-gradient operator to update only rotation matrices, preserving the integrity of the language modeling objective.
  • Significant reductions in perplexity were achieved on the LLaMA-2 7B model, demonstrating competitive performance against post-training quantization methods.
Read more
Local Stability and Gaussian Smoothing of Quantized Neural Networks
Sergey Salishev, Anton Makarov, Oleg Granichin
Theory Optimization Efficient ML
  • Gaussian averaging serves as a smooth surrogate for quantized neural networks, aiding in stability analysis.
  • The authors derive a local dimension-dependent estimate for the difference between quantized and smoothed models.
  • Closed-form Gaussian averages for ReLU and sign functions are computed and applied to a binary perceptron.
  • The concept of bounded local oscillation is introduced as a weak substitute for Lipschitz regularity in discontinuous settings.
Read more
From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python
Muntasir Adnan, Manile Srun, Carlos C. N. Kuhn
Large Language Models Reinforcement Learning Optimization
  • The ALPHA penalty is validated as an effective training signal for CWE prediction.
  • Delivery mechanisms significantly influence the effectiveness of the penalty as a training signal.
  • Reinforcement learning (GRPO) outperforms supervised methods under distribution shifts.
  • The best GRPO configuration achieves a substantial reduction in cumulative ALPHA penalty.
Read more
Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy Library
Cayan Deniz Kucuktopana, Javier Fumanal-Idocin, Richard Pitts, Javier Andreu-Perez
Interpretability
  • Introduces an interpretable regression extension for the Ex-Fuzzy library focused on Mamdani fuzzy inference.
  • Employs a target-aware partition initialization strategy using Fuzzy C-Means clustering to enhance interpretability.
  • Demonstrates that Gaussian partitions outperform trapezoidal partitions in terms of predictive accuracy and rule compactness.
  • Provides a transparent alternative to traditional black-box regression models, making it suitable for safety-critical applications.
Read more
Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines
Irena Girshovitz, Dan Zeltzer, Ran Gilad-Bachrach
Theory
  • Introduction of ARA framework for automated causal research.
  • Integration of protocol construction, synthetic data generation, and adversarial validation.
  • Focus on making silent failures in causal assumptions visible.
  • Evaluation showed ARA surfaces methodological concerns rather than just providing estimates.
Read more
The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks
Antonio Di Cecco
Theory Interpretability
  • Introduction of the quadrilateral loss as a differentiable penalty for measuring additivity in neural networks.
  • Demonstration of a strong accuracy-interaction trade-off, allowing for the removal of most interactions with minimal accuracy loss.
  • Surrender curves reveal that pre-regularization interaction magnitudes are unreliable predictors of retained interactions.
  • Comparison of various methods to achieve additivity, emphasizing the importance of behavioral constraints.
Read more
Detecting Neural Network Failures through Spectral Analysis of Internal Activations
Arunan J
Theory Interpretability Efficient ML
  • Introduction of 'Spectral Drift' as a failure signature in neural networks.
  • Development of the Self-Detecting Neural Networks (SDNN) framework for monitoring internal activations.
  • Significant performance improvement (79.0±25.3% AUROC) over confidence-based detection methods.
  • Use of curriculum learning to effectively train the detector on various failure scenarios.
Read more
Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignment Score For Retinal Disease Classification
Kritanu Chattopadhyay, Sayanjit Singha Roy, Soumya Chatterjee
Computer Vision Generative Models Interpretability
  • Introduction of CounterFundus, a CycleGAN-driven counterfactual explainability framework for retinal disease detection.
  • Development of the Counterfactual-Classifier Alignment Score (CCAS) to assess spatial agreement between counterfactuals and classifier saliency.
  • Demonstration of improved classification performance through CCAS-filtered counterfactual augmentations.
  • Provision of visually plausible disease-to-normal retinal translations for enhanced clinical interpretability.
Read more
Memoir: Should a Model Write to Its Memory While It Thinks?
Jaber Jaber, Osama Jaber
NLP Large Language Models Theory
  • Memoir architecture combines fast memory writing with latent recurrence for adaptive learning.
  • The coupled model shows a learning-speed penalty compared to the read-only model at fixed training steps.
  • Both models achieve the same performance ceiling after extended training, indicating no inherent capability disadvantage.
  • The predicted negative impact of memory rewriting on energy signals did not materialize.
Read more
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja, Saiyam Jain, Vaishnavi Shivkumar, Ponnurangam Kumaraguru
Large Language Models Optimization NLP
  • OPIUM mitigates the negative effects of activation steering, such as safety vulnerabilities and over-refusal.
  • The method employs dual-objective optimization to balance safety and utility in LLMs.
  • OPIUM is efficient, converging in minutes without the need for model weight updates.
  • The approach demonstrates that safety and utility can coexist in activation space.
Read more
On Optimization Complexity of Second-Order Certified Unlearning
Nikita Doikov, Anastasia Koloskova
Optimization Theory Efficient ML
  • Introduces a new second-order unlearning algorithm with state-of-the-art global convergence.
  • Establishes rigorous complexity bounds for certified unlearning and optimization accuracy.
  • Demonstrates that well-predicted removed data simplifies the optimization problem.
  • Proposes an anisotropic Gaussian mechanism aligned with the Hessian for effective unlearning.
Read more
AI-Driven Surrogate Models for Predicting Electrode-Scale Discharge Behavior in Lithium-Ion Batteries
Mengda Xing, Jean-Marie Lagniez, Alejandro Franco
Efficient ML Optimization Time Series
  • Introduction of an automated framework for integrating AI surrogates into legacy battery simulation software.
  • Development of a Swin3D Transformer model that processes volumetric mesh data for high-fidelity predictions.
  • Significant reduction in simulation time from hours to milliseconds while maintaining accuracy.
  • Enhanced spatial feature representation through Gaussian Positional Encoding.
Read more
Multi-modal transformer for signal classification in nanopore blockade experiments
Sandro Kuppel, Julian Hoßbach, Samuel Tovey, Christian Holm
Multimodal Time Series
  • Introduction of a multi-modal transformer architecture for nanopore signal classification.
  • Significant improvement in classification accuracy over existing methods.
  • Integration of multiple signal representations enhances feature extraction.
  • Attention analysis reveals complementary information from different modalities.
Read more
The C-index illusion: discrimination without calibration in published survival models
Rafael da Silva, Danilo Alvares
Theory
  • The C-index alone can mislead model evaluations by ignoring calibration and time-dependent accuracy.
  • Reproduced models showed high discrimination but significant calibration failures.
  • The study tested five hypotheses regarding the limitations of C-index evaluations.
  • The illusion of model performance is characterized by incorrect probability estimates rather than misranking.
Read more
StabilityBench: Benchmarking Instability in LLMs
Emma Kondrup, Zachary Yang, Anne Imouza, Reihaneh Rabbany
NLP Large Language Models
  • StabilityBench transforms single-turn benchmarks into multi-turn interaction histories to better simulate real-world AI behavior.
  • The framework incorporates user simulations and baiting techniques to assess model stability under varied conditions.
  • Evaluation of nine LLMs revealed significant performance instability, indicating limitations of traditional static benchmarks.
  • StabilityBench-Mini offers a cost-effective way to enhance evaluation realism without increasing resource requirements.
Read more
CruiseBench: A Real-Flight-Aligned N-CMAPSS Benchmark for Engine RUL Prediction
Pu Cheng, Qiang Miao
Time Series
  • Introduction of CruiseBench as a focused benchmark for cruise-stage RUL prediction.
  • Development of CPM-N-CMAPSS to isolate cruising intervals for better model evaluation.
  • Establishment of a fixed protocol for input features and evaluation metrics.
  • Demonstration of baseline model performance with TSMixer achieving the lowest RMSE.
Read more
Emergent Misalignment Recruits a Pre-existing Persona Subspace
Mohammed Suhail B Nadaf
NLP Large Language Models Theory
  • Fine-tuning on narrow bad advice can cause broad misalignment in language models.
  • A low-rank persona core shared across unrelated domains influences model behavior.
  • Projecting the persona subspace out during fine-tuning prevents misalignment.
  • Injecting the persona subspace into untouched models increases misalignment.
Read more
Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling
Kyunghoon Hur, Eunjung Jeon, Hyun Woo Kim, Gyubok Lee, Seongjun Yang
Time Series Theory Efficient ML
  • First empirical study comparing simulation pre-training and architecture-level priors for biokinetic knowledge injection in neural networks.
  • Both methods consistently improve performance over no-prior baselines across multiple datasets and microbial species.
  • Simulation pre-training is found to be the more effective and data-efficient method compared to architecture-level priors.
  • Practical guidelines for constructing simulation datasets are provided, highlighting the importance of parameter diversity.
Read more
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works
Yu Wang
Reinforcement Learning Large Language Models Theory
  • Dense prediction rewards can lead to catastrophic policy collapse in GRPO-trained LLM agents.
  • The 'dark room' pathology results from the amplification of within-group variance due to z-scoring normalization.
  • Removing standard deviation normalization can significantly improve training outcomes.
  • A variance-profile criterion is proposed to evaluate the safety of reward signals.
Read more
HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Multi-Section Spatial Gene Expression Prediction from H&E Histology
Kritanu Chattopadhyay, Soumya Chatterjee, Ondrej Krejcar, Debotosh Bhattacharjee
Graph Learning
  • Introduces Domain-Aware Edge-Weighted convolution for better tissue heterogeneity representation.
  • Achieves a mean PCC of 0.8314 on held-out genes, outperforming thirteen baselines.
  • Employs evidential uncertainty estimation with 90% empirical coverage for prediction reliability.
  • Fuses spot- and domain-level representations through Hierarchical DomainGCN and CrossScaleGate.
Read more
Regularized Optimization on Grassmann Manifold: Theory, Algorithm and Applications
Zhuan Liang, Zheng Zhai
Graph Learning Optimization Theory
  • Introduction of a novel RPMA framework for robust community detection and clustering.
  • Establishment of geometric properties and optimality conditions for rank-K projection matrices on the Grassmann manifold.
  • Development of efficient Riemannian optimization algorithms that leverage Cayley transformations.
  • Demonstration of superior performance of RPMA over traditional spectral methods in noisy conditions.
Read more
SCPP: A Unified Python Library for Soft Clustering
Kiyan Rezaee, Morteza Ziabakhsh, Artin Bahrampour, Seyed Mohammad Ghoreishi, Asal Khaje, Ali Sajedifar, Manny Chalak, Ava Zerafatangiz, Sadegh Eskandari
Optimization
  • SCPP provides a unified software abstraction for diverse soft clustering algorithms.
  • The library integrates 40 representative algorithms under a consistent API.
  • It includes a comprehensive benchmarking ecosystem for reproducible evaluation.
  • SCPP is designed for extensibility and seamless integration with existing scientific Python tools.
Read more
Scale-Aware Learning of Chaotic Dynamics on Unstructured Meshes via Binned Spectral Losses
Kanad Sen, Romit Maulik
Graph Learning Time Series Theory
  • Introduction of binned spectral losses for unstructured mesh surrogate modeling.
  • Utilization of graph-Laplacian frequency bands to replace traditional Fourier bands.
  • Development of scalable Chebyshev polynomial graph filters to avoid costly spectral decomposition.
  • Implementation of GLEAM for scale-aware supervision across graph hierarchies.
Read more
Zero-Shot Heart Rate Variability Forecasting from Consumer Wearables Using Time Series Foundation Models
Luukas Peräkylä, Fahad Sohrab, Ville Hautamäki, Merja Heinäniemi, Sui Huang, Pekka Abrahamsson
Time Series
  • Introduced a novel imputation method for handling fragmented HRV data from wearables.
  • Demonstrated that Time Series Foundation Models can outperform traditional forecasting methods without fine-tuning.
  • Achieved MASE values between 0.81 and 0.87, indicating strong forecasting performance.
  • Chronos and TimesFM emerged as the top-performing models for HRV forecasting.
Read more
Chronofy: A Temporal-Logical Decay Architecture for Information Validity in Time-Aware Retrieval-Augmented Generation
Muntaser Syed, Marius Silaghi, Sheikh Abujar, Sharun Akter
NLP Large Language Models Time Series
  • Chronofy integrates temporal validity into RAG systems through a three-layer architecture.
  • The framework employs learnable exponential decay functions for graph-based retrieval.
  • Signal Temporal Logic is applied to verify the temporal validity of knowledge rather than LLM output confidence.
  • Chronofy significantly improves retrieval precision and reduces temporal hallucination.
Read more
Beyond Heavy Log Curation: Perplexity-Based APT Detection via Unsupervised, Context-Augmented Language Models
Shoya Otsu, Kei Suzuki, Toshiaki Koike-Akino, Jing Liu, Ye Wang
NLP Large Language Models Time Series
  • CAPTAIN leverages pre-trained language models for APT detection with minimal preprocessing.
  • The model incorporates temporal context to enhance detection accuracy.
  • Smoothing filters are applied to perplexity scores for improved stability.
  • CAPTAIN shows competitive performance against existing methods while reducing engineering costs.
Read more
Local Causal Structure Learning in the Presence of Latent Variables and Selection Bias
Zheng Li, Hao Zhang, Ruxin Wang, Ruichu Cai, Kun Zhang, Feng Xie
Graph Learning Theory
  • Introduces LoCaLS, a local causal structure learning algorithm that handles latent variables and selection bias.
  • Establishes a theoretical connection between local and global causal structures.
  • Demonstrates higher structural accuracy and lower computational cost compared to existing methods.
  • Validates the approach with real-world gene expression data, yielding biologically meaningful results.
Read more
Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy
Jingyuan Li, Xiaoyi Jiang, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
Generative Models NLP Theory
  • Introduces Mean-to-Score (M2S) to ensure Bayes realizability in discrete diffusion models.
  • Identifies significant violations of Bayes realizability in existing SEDD models.
  • Demonstrates improved generative performance metrics over traditional models.
  • M2S constructs scores from predicted clean-token posteriors, enhancing model reliability.
Read more
Predicting Groundwater Arsenic Concentrations Using Graph Neural Networks
William Xing, Stephanie Yang, Aarush Bandemegal, Anushree Misra, Ananya Kalapatapu, Brennan Lagasse, Kevin Zhu
Graph Learning
  • Introduces a unified dataset of over 74,000 arsenic samples for improved prediction accuracy.
  • Demonstrates the effectiveness of graph neural networks in modeling spatial dependencies in arsenic concentrations.
  • Shows that GNNs can match or outperform traditional models like gradient-boosted trees in prediction tasks.
  • Addresses the limitations of previous studies that primarily focused on classification rather than regression.
Read more
Efficient Clustering with Provable Guardrails for LLM Inference at Scale
Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer
Large Language Models Efficient ML Optimization
  • Introduces a two-stage clustering algorithm that enforces similarity and attribute guardrails.
  • Achieves significant speed improvements (10–1000×) compared to standard clustering methods.
  • Successfully deployed on a large-scale application, reducing compute costs by 50-fold.
  • Scalable to tens of millions of samples while ensuring per-sample quality control.
Read more
Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts
Andrei Cristian Popescu, Haitz Sáez de Ocáriz Borde, Pietro Liò
NLP Large Language Models Efficient ML
  • Introduces a trajectory–readout perspective on adaptive depth in looped Transformers.
  • Demonstrates that fixed-prior depth supervision can create effective trajectories without input-dependent halting policies.
  • Shows that simple post-hoc confidence readouts can match or outperform learned gates in terms of compute-accuracy tradeoffs.
  • Identifies that issues in adaptive compute performance stem from the interaction between trajectory formation and exit selection.
Read more
External Clustering Validation by the Homogeneity-Parsimony Trade-off
Andreas Tiffeau-Mayer
Theory
  • Introduces normalized scores for clustering homogeneity and parsimony to evaluate clustering against known classes.
  • Establishes a mathematical foundation showing that these scores vary monotonically with cluster refinement.
  • Extends the information-theoretic framework to include set-matching and pair-based evaluation metrics.
  • Demonstrates the framework's applicability in feature selection and algorithm comparison.
Read more
Expanding Flow Maps
Sophia Tang, Pranam Chatterjee
Generative Models Graph Learning NLP
  • Introduction of Expanding Generative Flows (EFlows) for variable-dimensional generative modeling.
  • Development of Expanding Flow Maps (EFMs) that utilize an expand operator and a transport map.
  • Extension of the framework to discrete simplex space for variable-length discrete generation.
  • Demonstration of EFlows and EFMs in diverse applications including molecular generation and language modeling.
Read more
Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity
Hongnan Ma, Yiwei Shi, Mengyue Yang, Weiru Liu
Time Series
  • Existing time-series explanation methods often prioritize sufficiency, leading to the selection of non-essential subsequences.
  • TimePNS introduces a two-stage framework that incorporates counterfactual necessity into time-series explanations.
  • The framework learns a causal latent representation to better assess the necessity of temporal factors.
  • Experiments show that TimePNS outperforms strong baselines in identifying critical subsequences and reducing spurious explanations.
Read more
ELsAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
Mahdi Heidari, Mohammad Mahdi Rahimi, Jaekyun Moon
NLP Large Language Models Efficient ML
  • ELSAA combines sparse and low-rank attention approximations for improved efficiency.
  • The method introduces a denominator-aware fusion term to balance contributions from both branches.
  • It allows for longer-context training while preserving important token interactions.
  • The approach does not require materializing the full quadratic attention score matrix.
Read more
When Does Consensus Beat Voting? A Critical Analysis of Statistical Label Fusion in Medical Image Segmentation
Renjie He
Computer Vision
  • STAPLE does not provide a meaningful advantage over majority voting under typical conditions.
  • The algorithm is prone to instability and performance degradation under class imbalance.
  • Deep consensus methods with annotator embeddings show promise for improved segmentation.
  • Majority voting is recommended as the default method for fewer than 10 raters.
Read more
Geospatial Diffusion-based Evolution Synthesis (GeoDES) for Storm-Centered Weather Augmentation
Sonia Cromp, Satya Sai Srinath Namburi GNVV, Youran Wang, Grace Kisslinger, Frederic Sala, James Booth, Allegra LeGrande
Generative Models Time Series Theory
  • GeoDES synthesizes high-fidelity storm events using a storm-centered approach.
  • The model significantly outperforms existing weather prediction methods on key metrics.
  • GeoDES captures fine-scale storm dynamics without the smoothing typical of global models.
  • The architecture allows for computational efficiency, enabling training on standard consumer-grade GPUs.
Read more
Explanation-Based Runtime Verification for Trustworthy ML-driven Optical Networks
Omran Ayoub, Carlos Natalino, Ali Al Housseini, Felix Foschum, Philipp Morger, Tiziano Leidi, David Hock, Paolo Monti
Interpretability
  • Introduction of explanation-based runtime verification for ML decisions in optical networks.
  • Utilization of XAI techniques to assess the reliability of individual ML predictions.
  • Real-time evaluation of decision coherence and consistency with physical principles.
  • Demonstrated effectiveness in reducing false positive rates while preserving automation.
Read more
Joint Utilization of Geospatial and census proxies for Autoencoder-Assisted Downscaling (JUGAAD) of socioeconomic indicators in India
Aditya Dutt, Paul Gader, Aditya Singh
Computer Vision
  • Introduces JUGAAD, a deep learning framework for socioeconomic indicator downscaling.
  • Utilizes a three-step process involving data averaging, autoencoder compression, and regression modeling.
  • Validates predictions against ground-truth district-level NSSO indicators.
  • Demonstrates strong accuracy in predicting fine-scale socioeconomic indicators.
Read more
Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models
Dmitriy Poyarkov, Aleksei Staroverov, Aleksandr I. Panov
Reinforcement Learning Multimodal Robotics
  • Hybrid offline-online training can enhance the efficiency of RL in VLA models.
  • Incorporating offline supervision preserves OOD performance while reducing training costs.
  • Two guided PPO variants were proposed and evaluated, demonstrating improved performance over standard methods.
  • The study provides insights into the role of supervised priors in RL optimization.
Read more
Adversarial Frontiers: Minimum-Norm Attack Ensembles for Robustness Evaluation
Luca Scionis, Luca Melis, Maura Pintor, Fabio Brau, Ambra Demontis, Giorgio Fumera, Fabio Roli, Battista Biggio
Theory Optimization Computer Vision
  • Introduces a unified evaluation framework for adversarial robustness using minimum-norm attacks.
  • Defines attack and defense frontiers to provide worst-case robustness estimates and optimality rankings.
  • Establishes a controllable query budget for evaluating attack ensembles, allowing for flexible trade-offs.
  • Proposes the Defense Optimality Index (DOI) for stable, ε-independent defense rankings.
Read more
Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay
Xiaolong Li, Zhangchen Zhou, Zhi-Qin John Xu
Optimization Theory
  • Introduces weight-norm criticality as a mechanism for loss spikes in deep learning.
  • Demonstrates the interaction between normalization and weight decay as a destabilizing factor.
  • Establishes a theoretical framework linking weight decay to loss spikes through local sharpness of the loss landscape.
  • Empirically validates the proposed mechanism in Transformer and ResNet-50 architectures.
Read more
Generative Bayesian Filtering for State Estimation
Lei Cao, Sihang Feng, Jixin Yan, Tao Sun, Naichen Shi
Generative Models Time Series Theory
  • Introduction of Generative Bayesian Filtering (GBF) for improved state estimation.
  • Utilization of pretrained conditional generative models to enhance observation modeling.
  • Development of a novel backpropagation scheme for information transfer from high-dimensional observations to low-dimensional states.
  • Demonstrated improvements in accuracy and robustness in various applications, particularly under noisy conditions.
Read more
PhantomFill: When the Form Demands an Answer, Language Models Invent One
Rana Muhammad Usman
NLP Large Language Models
  • Language models are prone to hallucination when required to fill structured outputs, leading to fabricated answers.
  • The study introduces the Abstention-Affordance Ladder to analyze how output format affects model honesty.
  • PhantomFill benchmark measures Coerced Fabrication Rate and Escape Utilization Rate, addressing a gap in current evaluation metrics.
  • Models exhibit varying levels of resistance to fabrication based on their training and the specific output format used.
Read more
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
Yihan Wang, Zhong Guan, Haoran Sun, Jiale Huang, Likang Wu, Hongke Zhao
Reinforcement Learning Large Language Models Optimization
  • Introduction of Prefix-GRPO, a framework that enhances learning from teacher trajectories for small language models.
  • Decomposition of teacher trajectories into multiple training queries, allowing for differentiated learning.
  • Incorporation of a mechanism for optimizing historical tokens, improving the learning process.
  • Demonstrated improvements in agent performance across various environments compared to traditional methods.
Read more
Filter Learning for Subgraphs: Algebras and Performance Risk Bounds
Purui Zhang, Feng Ji, Yanan Zhao, Bihan Wen, Wee Peng Tay
Graph Learning Theory Optimization
  • Introduction of a systematic framework for subgraph filter learning (SFL).
  • Development of a distance-based subgraph filter algebra for effective approximation.
  • Establishment of finite-sample excess-risk analysis for ridge-regularized SFL.
  • Demonstration of improved performance over standard polynomial filtering methods.
Read more
HypNO: A Graph-Based Neural Operator with Physics-Informed Message Passing for Hyperbolic Conservation Laws
Dimitrije Ždrale, Cassie An Jeng, Katie Wang, Sonia Vanier, Alexandre Bayen, Hossein Nick Zinat Matin
Graph Learning
  • HypNO is a graph-based neural operator specifically for hyperbolic conservation laws.
  • It utilizes physics-informed message passing to respect physical properties like upwinding and entropy admissibility.
  • The model is benchmarked against established traffic-flow models, demonstrating superior performance in shock capturing.
  • HypNO offers a faster inference capability compared to traditional numerical methods after an upfront training phase.
Read more
Demographically-Informed Heat-Mortality Risk Curves via Risk Graph Neural Networks
Alex O. Davies, Eunice Lo, Rui Zhu
Graph Learning Time Series Interpretability
  • Introduces Risk Graph Neural Networks (RGNNs) to enhance DLNMs with demographic features.
  • Maintains interpretability of risk curves while improving predictive calibration.
  • Demonstrates effectiveness across multiple regions during extreme heat events.
  • Provides spatially-resolved mortality predictions relevant for public health policy.
Read more
ReliableTableQA: How Much Supervision Does Reliability Annotation Need?
Huei-Chung Hu, Hsin-Tai Wu, Koyo Kobayashi
NLP Large Language Models Reinforcement Learning
  • Introduction of a ten-category reliability taxonomy for tabular data queries.
  • Development of a data generation pipeline yielding 50,000 labeled examples.
  • Demonstration that minimal supervision (200 examples) can achieve high reliability scores.
  • GRPO provides limited benefits, primarily when SFT is under-trained.
Read more
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin
Multimodal Audio & Speech Large Language Models
  • Introduction of a three-tier symmetric dataset for audio reasoning.
  • Development of a cross-modal on-policy distillation framework (X3-OPD).
  • Dynamic alignment of reasoning trajectories between audio and text models.
  • Substantial performance improvements on audio and speech understanding tasks.
Read more
Bayesian uncertainty estimation improves clinical decision making in medical AI agents
Frederik Hauke, Patrick Wienholt, Christiane Kuhl, Dyke Ferber, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn
Computer Vision
  • Monte Carlo dropout provides reliable epistemic uncertainty estimates in medical AI models.
  • Integrating uncertainty signals improves error detection in clinical decision-making.
  • Presenting uncertainty as a binary flag is more effective than raw scores for clinical agents.
  • The study highlights the need for uncertainty representation in AI-driven medical diagnostics.
Read more
Exact ReLU realization of affine one-dimensional refinement iterates via residual memory and offset frames
Boldsaikhan Bolorkhuu, Tsogtgerel Gantumur
Theory
  • Introduces a residual memory controller for exact backward replay of states in ReLU networks.
  • Demonstrates that affine refinement iterates can be realized with fixed-width and linear depth.
  • Extends results to arbitrary compactly supported continuous piecewise linear forcing terms for M ≥ 3.
  • Provides a linear-depth realization for specific cases when M = 2.
Read more
KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers
Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann
Generative Models Efficient ML
  • KroQuant introduces a learned block-diagonal activation transform that is efficient and compatible with online constraints.
  • The method significantly reduces the parameter footprint compared to traditional dense invertible matrices.
  • KroQuant achieves faster inference times, with quantizer kernels running up to 14% faster than existing methods.
  • It demonstrates improved output quality on benchmark datasets compared to previous quantization techniques.
Read more
Autonomous Collaborative Learning Among an Ensemble of Tsetlin Machines with Consensus-Based Inference
Yehuda Rudin, Osnat Keren, Michal Yemini, Alexander Fish
Federated Learning Efficient ML Theory
  • Introduces a decentralized collaborative learning paradigm for Tsetlin Machines.
  • Utilizes consensus-based inference to integrate predictions from multiple agents.
  • Maintains data privacy by avoiding raw data exchange among agents.
  • Demonstrates comparable classification accuracy to centralized models.
Read more
Best-of-Evidence: Best-of-N Selection under Partial Verification
Cenwei Zhang, Teng Fang, Yuxia Wang, Derek Li, Bryan Dai, Lei You
Multimodal Optimization Theory
  • Introduces Best-of-Evidence (BoE) framework for selection under partial verification.
  • Utilizes a signed candidate-factor graph to represent claims and their relationships.
  • Establishes theoretical limits on evidence-driven improvements.
  • Demonstrates practical improvements in medical VQA tasks using BoE.
Read more
Agent-Centric Animal Pose Forecasting
Eyrun Eyjolfsdottir, Kristin Branson
Generative Models Robotics Time Series
  • Introduction of agent-centric autoregressive models for animal behavior forecasting.
  • Development of a library for designing and evaluating these models with biologically grounded metrics.
  • Successful modeling of social behavior in groups of courting Drosophila, capturing detailed motor patterns and behavioral motifs.
  • Framework allows for systematic comparison of model components and representations.
Read more
Thermodynamic Weight Decay: Exploring Grokking Acceleration via Attention Specific Heat
Chitraansh Pandey
Optimization Theory
  • Introduces CvAdamW, a thermodynamically-aware optimizer that adjusts weight decay based on attention specific heat.
  • Demonstrates that the peak of attention specific heat (Cv) reliably precedes the generalization transition in neural networks.
  • Identifies and resolves three failure modes in the training process, enhancing the reliability of the optimization.
  • Achieves significant reductions in grokking latency, improving generalization in a fixed compute budget.
Read more
CLOE: Christoffel Loss Autoencoder for Anomaly Detection
Léa Billet, Louise Travé-Massuyès, Elodie Chanthery, Alexandre Gaffet
Efficient ML Theory Optimization
  • CLOE integrates an autoencoder with a Christoffel Function-based anomaly detection method.
  • The method introduces a novel loss function that enhances representation learning for anomaly detection.
  • CLOE requires only one hyperparameter, reducing the complexity of model tuning.
  • The approach is computationally lightweight and designed for high-dimensional datasets.
Read more
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
Todd Y. Zhou
Reinforcement Learning Theory Optimization
  • Pass@k inversion occurs when RLVR improves pass@1 but reduces the number of distinct problems solved at larger k.
  • Boundary prompts, which contain rare correct trajectories, are particularly vulnerable to this inversion.
  • The paper introduces Per-Problem Base Anchoring (PBA) as a method to preserve rare correct trajectories during training.
  • Experiments show that PBA improves both pass@1 and high-budget coverage compared to matched controls.
Read more
M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data
Francesca Pia Panaccione, Carlo Sgaravatti, Marco Venere
Generative Models Multimodal Interpretability
  • M3-Gen integrates clinical metadata and histopathology images to generate gene expression profiles.
  • The framework utilizes contrastive learning to create a unified latent representation of multimodal data.
  • M3-Gen provides intrinsic explainability by identifying influential regions in histopathology images.
  • Evaluation on the TCGA dataset shows that the generated gene expression data is realistic and biologically coherent.
Read more
Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design
Tianming Han, Zhijie Pan, Wenchi Ge, Qi Zhao
Graph Learning Generative Models Optimization
  • Traditional SBDD is inadequate for complex diseases due to its focus on single targets.
  • Polypharmacology offers a more effective strategy by targeting multiple proteins simultaneously.
  • Geometric deep learning (GDL) provides a framework for overcoming limitations in multi-target drug design.
  • The paper reviews various GDL architectures and their applications in drug discovery.
Read more
Scaling Closed-Loop Feature Channel Configuration with LLMs
Tolgay Atinc Uzun, Radu Timofte, Dmitry Ignatov
Large Language Models Optimization Efficient ML
  • The study scales the closed-loop channel-configuration search to 250 candidates per cycle, significantly increasing the evaluation sample size.
  • A positive linear trend in mean accuracy is observed, with the best model achieving higher accuracy with fewer parameters.
  • Architectural regularities emerge from the larger sample, including non-standard channel widths and structured allocation patterns.
  • The findings confirm that LLM-driven channel search behavior is consistent across different sampling scales.
Read more
Post-Training in Time Series Foundation Models: A Unifying Framework
Shifeng Xie, Ambroise Odonnat, Zehao Xiao, Lei Zan, Malik Tiomoko, Lujia Pan, Themis Palpanas, Boris N. Oreshkin, Chenghao Liu, Keli Zhang
Time Series
  • Introduces a unifying framework for post-training methods in Time Series Foundation Models (TSFMs).
  • Categorizes post-training methods into five families, each addressing specific challenges in time series analysis.
  • Identifies limitations of current methods and suggests future research directions for improved model adaptation and reliability.
  • Highlights the importance of post-training for effective deployment of pretrained TSFMs in real-world applications.
Read more
The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability
Abigail Woodring, Adrian Chan, Rana Muhammad Shahroz Khan, Sukwon Yun, Chau-Wai Wong, Tianlong Chen
Large Language Models Efficient ML Theory
  • PortLLM demonstrates long-term temporal portability of LoRA patches across multiple continual pretraining steps.
  • The study empirically validates that repeated fine-tuning is not required for effective model adaptation.
  • Theoretical analyses reveal that near-orthogonality in high-dimensional spaces contributes to the observed performance of PortLLM.
  • The paper provides a geometric perspective on the loss landscape to facilitate comparisons of adaptation strategies.
Read more
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
Hao Liang, Qifeng Cai, Yibo Lin, Jianzhuo Du, Qifeng Xia, Sizhe Qiu, Linzhuang Sun, Meiyi Qiang, Zhaoyang Han, Xiaochen Ma, Bohan Zeng, Ruichuan An, Conghui He, Wentao Zhang
NLP Large Language Models
  • DataPrep-Bench is the first unified benchmark for evaluating LLMs in data preparation tasks.
  • It assesses both data construction and data quality evaluation capabilities under a shared protocol.
  • The introduction of the Distributional Alignment Score (DAS) provides a robust metric for evaluating dataset quality.
  • Data-Construction-Skill significantly enhances data construction performance, outperforming existing methods.
Read more