AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

57 Papers today
8h Update frequency
7 Days of history
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
Mobina Kashaniyan, Ali Jannesari
Large Language Models Efficient ML NLP
  • Candidate count alone does not adequately describe the system costs of multi-candidate inference.
  • Increasing the number of candidates improves accuracy but also increases energy consumption and latency.
  • Batched generation calls are significantly more efficient than serial execution of candidates.
  • The paper provides recommendations for reporting practices in multi-candidate inference studies.
Read more
Learning-Induced Dynamical Transition in Recurrent Neural Networks
Varun Vaidya
Theory
  • Learning in RNNs can induce a transition from chaotic to stable dynamics.
  • A non-equilibrium dynamical mean-field theory (DMFT) is developed to analyze this transition.
  • The study identifies critical parameters that separate chaotic and stable regimes during learning.
  • The theory predicts the evolution of network outputs and shows quantitative agreement with simulations.
Read more
QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization
Yujie Li, Zezhi Shao, Chengqing Yu, Yisong Fu, Weijie Zhu, Yifan Du, Jilin Hu, Bin Yang, Yongjun Xu, Fei Wang
Time Series Efficient ML Optimization
  • QUALS addresses data diversity issues in time series forecasting.
  • The framework includes pattern quantization and learnability synchronization mechanisms.
  • Models trained on QUALS achieve superior zero-shot performance with less data.
  • The approach effectively manages skewed pattern distributions and learnability discrepancies.
Read more
How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?
Nguyen Dung Son, Dang Quang Minh, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh
NLP Large Language Models
  • Rubric-decomposed prompting outperforms holistic prompting in most cases.
  • The mapping of trait scores to prompt scores is fragile and requires careful calibration.
  • Essay length impacts scoring accuracy, with smaller models showing variability in performance.
  • The best local model configuration achieved a QWK of 0.388, below human scoring standards.
Read more
Radio-Frequency Convolutional Neural Networks
Zhihui Gao, Shi-Yuan Ma, Yiran Chen, Dirk Englund, Tingjun Chen
Efficient ML
  • RF-CNN repurposes existing wireless communication hardware for efficient CNN inference.
  • The approach achieves low energy consumption, down to 0.72 femtojoules per multiply-accumulate operation.
  • RF-CNN can handle deep CNNs with millions of parameters while maintaining performance close to full precision.
  • The method utilizes the native operations of frequency mixers, avoiding the need for additional dedicated hardware.
Read more
Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map
Patricia Medina, Hy P. G. Lam
Theory
  • Establishes sharp reconstruction bounds for autoencoders using the same forward map.
  • Derives a reconstruction-derivative error bound based on Jacobian singular values.
  • Demonstrates that affine maps can achieve the derived bounds at any depth.
  • Validates theoretical predictions with empirical results from a large LiDAR dataset.
Read more
MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards
Shihao Liu, Hao Yin, Lijun Liu, Zhengzong Chen, Yuanyuan Zhao, Fei Huang
NLP Large Language Models Reinforcement Learning
  • Identifies and addresses scheduling and reward gaps in RL-based tool learning.
  • Introduces Model-Aware Curriculum Learning (MACL) for adaptive sample difficulty.
  • Presents Hierarchical Tool-call Gated Reward (HTGR) for structured reward evaluation.
  • Demonstrates superior performance of MATCH over existing baselines in tool learning tasks.
Read more
ZeroHAT: Behavior-Conditioned Zero-Shot Human Activity Trace Generation
Rongchao Xu, Dahai Yu, Lin Jiang, Guang Wang
Generative Models Time Series Optimization
  • ZeroHAT generates synthetic HATs without requiring target-region data, using only source-region data and contextual information.
  • The framework includes innovative components for intent extraction, behavioral cloning, and activity realization.
  • ZeroHAT significantly outperforms existing methods in terms of utility and fidelity across multiple target regions.
  • The approach demonstrates the effectiveness of transferring behavioral patterns across regions with different POI distributions.
Read more
Expected Hypervolume Maximization for Multiobjective Optimization under Uncertainties
Victor Trappler
Optimization
  • Framing multiobjective optimization under uncertainty as a Bayesian decision problem.
  • Maximizing the expected hypervolume as a utility function for optimization.
  • Utilizing Gaussian Processes as surrogate models to reduce computational costs.
  • Introducing active learning strategies to improve surrogate model performance.
Read more
Local Sparsity Enables Unsupervised LLM Safety Detection
Xin Chen, Gil Kur, Alexander Shevchenko, Andreas Krause
Large Language Models NLP Theory
  • Proposes an unsupervised anomaly detection framework for LLM safety that does not rely on unsafe training data.
  • Utilizes the Linear Representation Hypothesis to exploit local sparsity in LLM activations for improved anomaly detection.
  • Demonstrates that locally sparse methods can achieve near-optimal performance with minimal computational resources.
  • Validates the approach across multiple LLM architectures and safety-specific datasets.
Read more
Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance
Girish Keshav Palshikar
Theory Optimization Efficient ML
  • Introduces the notion of allied datasets with similar class labels but heterogeneous feature spaces.
  • Proposes a method for merging feature spaces using matrix completion techniques.
  • Demonstrates that classifiers trained on unified datasets outperform those trained on individual datasets.
  • Highlights the potential for improved generalization and knowledge transfer across datasets.
Read more
AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection
Omran Berjawi, Walid Fahs, Rida Khatoun
Multimodal
  • Introduction of AURA, a multi-modal email threat detection system.
  • Utilizes an Adaptive Uncertainty Router to manage prediction uncertainty.
  • Achieves high performance on both in-distribution and out-of-distribution datasets.
  • Combines URL analysis with semantic content analysis for improved detection.
Read more
Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation
Jing Zhang
Reinforcement Learning Robotics Theory
  • RSIQL improves offline GCRL by providing additional reward signals at intermediate states.
  • The method does not require a hierarchical policy, simplifying the learning process.
  • Experiments show significant performance improvements over existing methods.
  • The approach addresses the issue of delayed goal-completion supervision effectively.
Read more
Seismic Site Response Prediction from Sparse Observations Using Finite-Element-Pretrained Latent Dynamics
Yi Zhu, Su Chen, Xiaojun Li
Time Series
  • FLARE-T improves seismic site response predictions by integrating finite-element simulations with sparse field observations.
  • The framework learns low-dimensional latent dynamics to connect input and output accelerations at multiple depths.
  • Validation results show reduced prediction errors in acceleration histories and response spectra compared to conventional models.
  • The approach demonstrates robustness across different earthquake intensities and source models.
Read more
PhyRestore: Physics-Structured Latent-Factor Restoration
Ahmed Shafee, Chayan Lahiri
Time Series Theory Optimization
  • PhyRestore restores corrupted physical factors from bitemporal spatial observations.
  • The framework reconstructs temporal changes through the known RUSLE relationship.
  • Factor restoration improves recovery of rare high-magnitude changes under certain conditions.
  • The study compares multiple learning pathways for soil-loss prediction.
Read more
Online Adaptive Kernel Mixing for Gaussian Process Decision Making
Kavin Aravindan, Mani Tej Sriram, Gautam Dasarathy, Tejas Bodas
Optimization Theory
  • HACK GPs adaptively select kernels in Gaussian Processes to improve decision-making performance.
  • The method treats kernel selection as an online learning problem using AdaHedge for dynamic updates.
  • Two variants of HACK are introduced: Mixture of Gaussians and categorical sampling.
  • The approach shows improved performance over standard kernels and ensemble methods in empirical evaluations.
Read more
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan, Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah
Reinforcement Learning Large Language Models
  • ActObs introduces a new SFT approach that includes observation prediction, enhancing agent exploration in RL.
  • The method shows superior performance in pass rates and task diversity compared to traditional action-only SFT.
  • Observation supervision helps maintain policy entropy, facilitating better exploration and learning dynamics.
  • The findings suggest that SFT objectives significantly influence downstream learning and exploration capabilities.
Read more
Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not
Rubén Darío Guerrero
Optimization Theory
  • Introduction of Stiefel Attention, optimizing transformer projection matrices on the Stiefel manifold.
  • Development of a Riemannian Adam optimizer that respects the geometry of the manifold.
  • Significant performance improvements on benchmarks, with validation accuracy rising from 61.1% to 97.0%.
  • The method's advantages grow with increasing data, indicating a strong relationship between geometry and optimization.
Read more
Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization
Donney Fan, Colin Doumont, Aleksandra Kalisz, Paul Duckworth, Jacob R. Gardner, Henry Moss, Geoff Pleiss
Generative Models Optimization Efficient ML
  • Introduces a new LSBO algorithm that significantly reduces computational overhead.
  • Utilizes linear surrogates constrained to a spherical domain for efficient optimization.
  • Achieves over 100× speedup compared to existing BO methods while maintaining high sample efficiency.
  • Demonstrates effectiveness across molecular design and image generation tasks.
Read more
When Does Retrieval Help Time-Series Forecasting?
Mert Onur Cakiroglu, Elham Buxton, Mehmet Dalkilic, Hasan Kurban
Time Series
  • Retrieval benefits in time-series forecasting are primarily determined by the ratio of lookback window length to dominant seasonal period.
  • A simple control method can outperform complex retrieval plug-ins under certain conditions, particularly when the seasonal structure is strong.
  • The study introduces a regime map and two statistics to predict the effectiveness of retrieval mechanisms before deployment.
  • Retrieval mechanisms are less effective when the training data lacks a concentrated seasonal structure.
Read more
Recursive Quantum Long Short-Term Memory for Stable Short-Horizon Temperature Forecasting
Mu-En Lee, Yen-Ku Liu, Samuel Yen-Chi Chen, Yun-Cheng Tsai
Time Series
  • Introduction of Recursive Quantum Long Short-Term Memory (RQLSTM) architecture.
  • Empirical evaluation against standard QLSTM for temperature forecasting.
  • RQLSTM shows improved convergence and predictive accuracy.
  • Lower mean absolute error and root mean squared error compared to QLSTM.
Read more
Bayesian Optimization with Rich Auxiliary Information via LLMs
Tejus Gupta, Efe Mert Karagözlü, Rohit Sonker, Barnabás Póczos, Jeff Schnieder
Optimization Large Language Models
  • LLMs can effectively leverage rich auxiliary information to enhance Bayesian Optimization.
  • Two novel methods are proposed for incorporating LLM-derived priors into BO.
  • Utilizing auxiliary observations beyond standard (x, y) interactions improves optimization performance.
  • The methods outperform both traditional BO and existing LLM-based optimization techniques.
Read more
Layer-wise Curriculum Learning for Efficient LLM Compression
Donggeon Lee, Dooyeon Na, Seungmin Oh, Jongbin Ryu
Large Language Models Efficient ML Optimization
  • Introduces layer-wise curriculum learning for efficient LLM compression.
  • Addresses cumulative error phenomenon through theoretical analysis and practical implementation.
  • Utilizes feature caching and multi-threading to enhance computational efficiency.
  • Achieves over 50% reduction in GPU memory usage and training hours compared to traditional methods.
Read more
COMPASS: Ordered Clustered Routing at 100K Scale
Ido Greenberg, Hugo Linsenmaier, Piotr Sielski, Shie Mannor, Alex Fender, Gal Chechik, Eli Meirom
Optimization
  • COMPASS is a globally-coordinated algorithm that optimizes the OCTSP using raw distance inputs.
  • The algorithm can scale to 100K synthetic nodes and 28.5K real-world e-commerce nodes, achieving state-of-the-art results.
  • COMPASS avoids a quality ceiling by modeling global dependencies between clusters.
  • The paper introduces a new benchmark suite for large-scale OCTSP, facilitating future research.
Read more
Dynamic Generalized Gromov-Wasserstein Optimal Transport
Junda Ying, Zhiwei Zeng, Peijie Zhou, Lei Zhang
Optimization Theory Time Series
  • TP-DATE provides a dynamic generalization of Gromov-Wasserstein optimal transport.
  • The framework allows for structure-aware transport in applications like spatial transcriptomics.
  • Travelling-pair flow matching facilitates the modeling of interacting conditional paths.
  • TP-DATE outperforms existing static methods in preserving spatial structure and reconstructing dynamics.
Read more
COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression
Zewen Yang, Xiaobing Dai, Zhenxiao Yin, Hang Zhao, Zhijun Li, C.C. Chan
Robotics Theory Efficient ML
  • Introduction of COIN-GP framework for cooperative online learning in distributed systems.
  • Utilization of Gaussian Process regression to handle partial measurements and unknown dynamics.
  • Development of a novel data collection strategy with theoretical guarantees.
  • Derivation of error bounds for state and model estimation.
Read more
Evaluating Explanation Methods by the Predictors They Induce
Jacob Selbæk, Hugo L. Hammer
Interpretability Theory
  • Introduces a framework to evaluate explanation methods based on their predictive power.
  • Demonstrates that the effectiveness of explanations varies with feature dependence.
  • Finds that existing quality metrics can fail to accurately rank explanation methods.
  • Provides a controlled study across 13 real datasets and 9 synthetic designs.
Read more
Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems
Mohammed El Hanjri, Anas Abouaomar, Hamidou Tembine, Abdellatif Kobbane
Federated Learning Time Series
  • Introduces a coalition formation approach for federated learning that adapts to client heterogeneity.
  • Utilizes opinion dynamics to form coalitions based on local model weights, enhancing model compatibility.
  • Demonstrates significant improvements in forecasting accuracy for water consumption compared to traditional methods.
  • Achieves coalition structures with no additional client-side computation or communication overhead.
Read more
Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection
Rishi Bharadwaj, Yadati Narahari
Reinforcement Learning Optimization Theory
  • The paper highlights the exclusion of smallholder farmers from carbon farming programs due to contract design issues.
  • A novel framework is proposed that combines moral hazard and adverse selection in a multi-season contract design.
  • The Realised Adoption Share (RAS) metric is introduced to assess the impact of MRV costs and profit maximization on farmer participation.
  • The study finds that profit-maximizing contracts disproportionately benefit larger farms, exacerbating smallholder exclusion.
Read more
CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning
Xiang Zou, Shengzhu Shi, Junqi Gao, Zhichang Guo
Reinforcement Learning Robotics Theory
  • Introduction of CARE-VI, a framework for improving target reliability in off-policy actor-critic learning.
  • Development of three components: CARS, SEVA, and DARE, which address candidate selection, value assessment, and residual regulation.
  • Theoretical analysis provides error bounds for the proposed methods, ensuring fixed-policy recovery.
  • Empirical results show CARE-VI achieves the highest mean return across multiple tasks and configurations.
Read more
Distributionally Robust Federated Learning with Multi-Source Data
Yingzhu Liu, Zhongkui Li, Pengcheng You, Ashish Cherukuri
Federated Learning Optimization Theory
  • Introduces a global ambiguity set that captures both cross-client mixture uncertainty and within-client distributional ambiguity.
  • Establishes out-of-sample performance guarantees for the proposed robust solution.
  • Develops a federated algorithm (BiDRO-FL) with proven convergence under relaxed conditions.
  • Validates the effectiveness of the proposed framework through simulations.
Read more
Labeled Incidence Structures for Native Transformer Modeling of Text, Knowledge Graphs, and Hypergraphs
Mahesh Godavarti
NLP Graph Learning Theory
  • LIS enables a single transformer to process text, KGs, and hypergraphs natively without architecture changes.
  • The multiplicative addressing method preserves joint information between roles and relations, unlike additive methods.
  • Content-computed operators address storage order issues in knowledge repositories, improving model robustness.
  • The framework shows potential for unified training across different structured data types.
Read more
A Policy Profile for Croissant: Refusal as a Property of the Dataset
Alexander Chernov
Theory
  • Introduces an additive policy profile for Croissant with specified evaluation semantics.
  • Defines a closed operator language for dataset operations and conditions.
  • Demonstrates that the profile's decisions align with existing descriptor records.
  • Presents a minimal evaluation cost for implementing the profile.
Read more
SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems
Babak Sarani, Rahman Ardakanian, Ali Mousavi
Optimization Interpretability Theory
  • Introduction of SoftTri, a differentiable triangular membership function that enhances optimization stability.
  • Closed-form analytical gradients derived for efficient backpropagation in training.
  • SoftTri maintains the interpretability and locality of classical triangular MFs while providing smoothness.
  • Experimental results show improved performance over classical triangular MFs and comparable results to Gaussian MFs.
Read more
Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort
Antony Garcia, Gabrielle Britton, Alcibiades Villarreal, Diana Oviedo, Giselle Rangel, Xinming Huang
Interpretability
  • I-309 (CCL1) was identified as a significant predictor of cognitive impairment, with a notable increase in predictive accuracy.
  • The study employed a leakage-safe machine learning methodology to ensure robust results.
  • The findings underscore the need for external validation of biomarkers in predicting cognitive impairment.
  • The research distinguishes between statistical significance and predictive utility in the context of small clinical datasets.
Read more
Graph-Based Stochastic Power-UCT: Monte-Carlo Graph Search with Power Mean Estimation
Tung Tran, Viet Bao Mai, Hoang Ta, Tuan Dam
Reinforcement Learning Graph Learning Theory
  • GS-Power-UCT merges states at the same depth to reduce redundancy in simulations.
  • The algorithm maintains separate values for states accessed at different depths, preserving estimation accuracy.
  • GS-Power-UCT achieves a convergence rate of O(n−1/2) for root estimates, matching traditional methods.
  • Two variants of the algorithm are proposed, with GS-Power-UCT-F+ showing improved performance through adaptive horizon control.
Read more
Amortizing Physics-Informed Neural Solvers via Graph Hypernetworks
Cheng Jing, Abhishek Verma, Kallol Bera, Yixuan He, Kookjin Lee
Graph Learning Optimization Theory
  • Introduces operator-graph conditioning for amortized PINNs.
  • Demonstrates improved accuracy in high-reaction scenarios using term-based descriptors.
  • Shows that graph conditioning yields lower mean error on unseen coupling compositions.
  • Identifies scenarios where traditional coefficient conditioning performs best.
Read more
Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
Zhilong Zheng, Letian Tao, Yang Guan, Yujie Yang, Wei Xiong, Kehua Sheng, Bo Zhang, Jingliang Duan, Keqiang Li, Shengbo Eben Li
Theory Efficient ML Large Language Models
  • Introduces Parameter Space Orthogonality as a feasible condition for preventing catastrophic forgetting.
  • Develops a post-hoc, tuning-agnostic framework (JANUS) that rectifies parameter updates after fine-tuning.
  • Employs a Multi-step Adaptive Rectification mechanism to dynamically adjust parameter updates.
  • Demonstrates significant improvements in knowledge recovery with minimal impact on task performance.
Read more
A Noise Optimum in Rehearsal-Free Continual Learning: Isolation, Mechanism, and Scope
Gunner Levi Howe
Theory
  • Injecting stochastic noise into consolidation rules can improve task retention up to an optimal level.
  • The optimal noise level is achieved through coherent restoring forces directed towards consolidated weights.
  • The coupling of anchor gain to noise variance is a critical factor in achieving retention improvements.
  • The noise optimum is dependent on shared task structure and is absent in tasks with permuted structures.
Read more
Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems
Xianjian Xie, Hao Yan
Federated Learning
  • Introduces a hierarchical decomposition for modeling heterogeneous federated clients.
  • Employs privacy-preserving federated variational inference to keep raw data local.
  • Supports uncertainty-aware predictions, crucial for risk-sensitive applications.
  • Demonstrates effectiveness in real-world applications like fault classification and air-quality modeling.
Read more
SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption
Wentao Zhang, Yifan Zhu, Yutong Zhang, Wentao Mo
Multimodal
  • Identifies a granularity mismatch in existing batch-level gradient modulation under heterogeneous multimodal corruption.
  • Proposes SAGG, which utilizes sample-level gating for unbiased gradient estimation.
  • Demonstrates that SAGG-based SGD converges at a rate of O(1/√T) without residual bias from corruption.
  • Provides a certified robustness framework for evaluating multimodal models.
Read more
Enhanced Agriculture-informed Neural Network by Domain Knowledge
Ci Lin, Futong Li, Rose Chong-Wu, Tet Yeap, Iluju Kiringa
Interpretability
  • KAINN integrates domain knowledge into deep learning for improved N2O emission predictions.
  • The framework incorporates key environmental processes to enhance model interpretability.
  • KAINN outperforms traditional neural network models and the original AINN in accuracy.
  • The approach demonstrates reduced uncertainty and improved stability in parameter trajectories.
Read more
Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements
Caterina Amendola, Giulia Maffeis, Lorenzo Buffoni, Lorenzo Chicchi, Francesco Coghi, Duccio Fanelli, Raffaele Marino, Fabrizio Martelli, Riccardo Paoli, Lorenzo Pattelli, Lorenzo Spinelli
Theory Optimization Efficient ML
  • Proposes a machine learning framework for reconstructing optical properties in bilayered media.
  • Generates a robust synthetic dataset using Monte Carlo simulations for training.
  • Achieves higher accuracy and faster reconstruction compared to traditional model-based methods.
  • Estimates parameter space dimensionality without prior information about layer count.
Read more
Towards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment
Stefanos Gkikas, Christian Arzate Cruz, Calvin Joseph, Giorgos Giannakakis, Raul Fernandez Rojas
Multimodal
  • Introduces a modality-agnostic framework for cognitive workload assessment using a hierarchical Transformer architecture.
  • Evaluates all combinations of five biosignal modalities in a systematic manner.
  • Finds EEG to be the most effective single modality for cognitive workload assessment.
  • Demonstrates that combining all five modalities achieves the highest performance scores.
Read more
One Intervention per Component is Enough: Towards Identifiability in Linear Stochastic Dynamics from Steady State
Saber Salehkaleybar
Theory Time Series Optimization
  • One intervention per strongly connected component (SCC) is sufficient for parameter recovery of the OU process.
  • A recursive algorithm is developed to topologically order SCCs and isolate marginal dynamics.
  • The proposed regularized least-squares estimator effectively minimizes residuals from steady-state equations.
  • Theoretical results are validated through empirical experiments, confirming the approach's effectiveness.
Read more
Subdomain-aware representation compression for pretrained image embeddings
Poowanut Niamluang, Jittat Fakcharoenphol
Computer Vision Efficient ML
  • Dimensionality reduction techniques can be effectively tailored for subdomain representation compression.
  • PCA and LDA show significant improvements in both space efficiency and accuracy for pretrained image embeddings.
  • The approach allows for effective transfer learning capabilities using compressed representations.
  • Experiments demonstrate that compression can outperform traditional full-embedding methods.
Read more
Radio Frequency Detection and Classification of Microplastics in Water
Jaden Tolbert, Md Saiful Islam, Pingshan Wang
Theory Efficient ML Multimodal
  • Introduction of a novel RF dielectric spectroscopic cytometry platform for microplastic detection.
  • Achieved high classification performance for eight microplastic classes in deionized water.
  • Demonstrated robustness of the method in saline environments.
  • Utilized machine learning to analyze RF scattering parameters for material classification.
Read more
Score Centering Stabilizes Off-policy Reinforcement Learning
Martin Marek, Max Ryabinin
Reinforcement Learning Large Language Models Efficient ML
  • Score centering effectively cancels drift caused by training-inference mismatch (TIM).
  • It outperforms traditional importance sampling methods under severe quantization conditions.
  • The method is additive and can be combined with importance sampling for improved performance.
  • The approach allows for stable training of large language models while maximizing hardware utilization.
Read more
Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning
Simon Süwer, Julian Klemm, Elisa Acitelli, Mathieu Almeida, Lucia Altucci, Zsolt Bagyura, Michelangela Barbieri, Zsolt-Zoltán Bedő, Rosaria Benedetti, Béla Bihari, Csongor Csalóka, Lucia Dicunta, Stanislav Ehrlich, Bjoern M. Eskofier, Sándor-József Fejér, Georg Fröwis, Walter Hötzendorfer, Alexandra Kautzky-Willer, Jens Johann Georg Lohmann, Marianna Maranghi, Lorenzo Marconi, Rudolf Mayer, Wouter Leonard Megchelenbrink, Monika Moga, Adham Mottalib, Sanjeev Mehta, Madeleine Müller, Thomas Nyström, Balázs-Attila Orbán, Paul O'Toole, Giuseppe Paolisso, Paolo Parini, Matteo Pedrelli, Enrico Petrillo, Philipp Poindl, Niklas Probul, Anastasia Pustozerova, Tanja Šarčević, Lukas Weilguny, Jan Baumbach, Andreas Maier
Federated Learning
  • FL-Net is a comprehensive federated learning framework tailored for multi-center medical research.
  • The framework fulfills five essential requirements for federated clinical research that existing frameworks do not meet.
  • FL-Net enables the reuse of harmonized data and workflows, promoting collaboration while ensuring data privacy.
  • The framework supports up to 50 concurrent clients in federated workflows, showcasing its scalability.
Read more
Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning
Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Yian Ma, Lianhui Qin
Multimodal Robotics NLP
  • Uni-LaDiR introduces a unified latent space for multimodal reasoning, improving the integration of information from different modalities.
  • The framework employs a unified encoder to generate shared thought tokens, enhancing the reasoning process by reducing modality-specific representation issues.
  • Diffusion is used to predict the next reasoning steps, allowing for flexibility in generating multiple valid outputs.
  • Joint training of the encoder and diffusion model leads to better task-relevant thought token generation.
Read more
Generative Query Suggestion via Intent Coverage and Query-Level Credit Assignment
Xinpeng Liu, Lu Ma, Jiayi Qiao, Mengyu Zhou, Linglong Li, Xiaofeng Bian, Haonan Chen, Xiaoxi Jiang, Guanjun Jiang
NLP Reinforcement Learning Generative Models
  • Introduces an Intent-Driven Query Suggestion Framework with dual-stage optimization.
  • Addresses intent-aware diversity modeling and query-level credit assignment challenges.
  • Implements an Intent-Aware Diversity Reward to optimize intent coverage.
  • Demonstrates significant improvements in user engagement metrics through extensive testing.
Read more
PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers
Jiachen Yao, Zi-Siang Hsu, Xi Deng, Aditi Gupta, Xin Ju, Sally M Benson, Gege Wen, Anima Anandkumar
Generative Models
  • PosteriorBench provides a comprehensive framework for evaluating generative inverse solvers beyond point estimates.
  • The benchmark includes four diverse scientific inverse problems, enabling a broad assessment of solver performance.
  • A five-metric evaluation suite allows for detailed analysis of distributional accuracy and uncertainty quantification.
  • Neural operators demonstrate improved robustness in resolution, addressing gaps in current solver performance.
Read more
Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data
Kaihua Ding
Large Language Models
  • Trained classical models outperform frozen LLMs in 86% of evaluated cases with minimal labeled data.
  • The crossover point where classical models become preferable is around 6% of the training set.
  • Few-shot prompting does not yield consistent performance improvements for LLMs.
  • The performance of LLMs is sensitive to the semantics of input features.
Read more
Smart Insole Human Activity Recognition for Continuous Monitoring in Elderly Care
Edwin Rios, Antony Garcia, Fengpei Yuan, Xinming Huang
Robotics Time Series Multimodal
  • Development of a smart insole platform for continuous monitoring of elderly mobility states.
  • Integration of pressure sensors and IMU for comprehensive activity recognition.
  • High performance of Histogram-Based Gradient Boosting in classifying mobility states.
  • Demonstration of effective participant-independent validation for model robustness.
Read more
Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training
Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang, Parth Shroff, Ishan S. Khare, Hermann Kumbong, Azalia Mirhoseini
NLP Large Language Models Efficient ML
  • Introduction of Block Parallelism (BP) for BDLM training, reducing cross-rank communication.
  • Development of Context-Sharded Block Parallelism (CSBP) to efficiently handle long-context training.
  • Significant throughput improvements of 1.18–1.45× for supervised fine-tuning and 1.27–1.33× for autoregressive to BDLM conversion.
  • CSBP accelerates speculative decoding training by up to 7.59× at 1M context length.
Read more
SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting
Abraham Ezema, Chijioke Eze, Ferdinanda Ponci, Antonello Monti
Time Series
  • SETTer introduces decoupled self-attention and hybrid masking techniques for improved forecasting.
  • The model effectively captures both short- and long-term dependencies in multivariate time series data.
  • SETTer provides explainable structures to enhance interpretability of its predictions.
  • The model outperforms existing state-of-the-art methods in 88% of scenarios tested.
Read more
Compressed Active Subspaces for Scalable Bayesian Inference
Thomas Flynn, Sanket Jantre, Byung-Jun Yoon, Kibaek Kim
Efficient ML Theory Optimization
  • Introduction of Compressed Active Subspaces (CAS) for scalable Bayesian inference.
  • CAS reduces memory requirements for active subspace construction in high-dimensional models.
  • Demonstrated scalability on neural networks while maintaining predictive performance.
  • Enables robust uncertainty quantification in large models.
Read more