AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

63 Papers today
8h Update frequency
7 Days of history
TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series
Nicolas Zumarraga, Lorenzo Steno, Ning Wang, Max Rosenblattl, Thomas Kaar, Maxwell A. Xu, Kevin O'Sullivan, Markus Kreft, Elgar Fleisch, Paul Schmiedmayer, Patrick Langer, Robert Jakob
Time Series NLP Reinforcement Learning
  • TimeRLM utilizes Recursive Language Models to improve anomaly localization in long-context time-series data.
  • The introduction of the AnomalyXL benchmark allows for precise evaluation of anomaly detection capabilities.
  • TimeRLM outperforms existing models, achieving significant improvements in localization and classification tasks.
  • Reinforcement learning further enhances the model's performance and efficiency.
Read more
Amortized Interventional Forecasting for Multivariate CIR Processes
Andreas Sauter, Sumit Sourabh, Drona Kandhai, Erman Acar
Time Series
  • Introduces CIR-ACTIVA, an amortized model for causal effect estimation in financial time series.
  • Develops a causal multivariate CIR data-generating process to provide necessary ground truth for evaluation.
  • Demonstrates superior performance of CIR-ACTIVA over traditional observational forecasting methods.
  • Enables the analysis of interventional scenarios, such as stress testing in financial markets.
Read more
Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation
Seyed Kahaki, Shijie Li, Weijie Chen, Nicholas Petrick
Generative Models Computer Vision
  • Introduction of a two-step training approach for denoising diffusion models.
  • Development of domain-specific modified FID and IS metrics using pathology-pretrained foundation models.
  • Empirical validation of the correlation between synthetic data quality metrics and downstream task performance.
  • Demonstration that increasing data variety is more beneficial for model performance than improving individual image quality.
Read more
Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL
Yi Yang, Zhennan Chen, Mingfeng Lv, Hanlei Li, Zhengsen Ruan, Lvqing Yang
Reinforcement Learning Robotics Theory
  • CSDG separates local corrections from in-sample targets to control OOD value contributions.
  • The method derives theoretical bounds for correction propagation and performance evaluation.
  • Empirical results show improved performance in offline RL tasks without pessimistic penalties.
Read more
WorldDynCache: Risk-Controlled Latent Dynamics Approximation for Diffusion World Model
Leyang Chen, Junyi Wu, Shaoqiu Zhang, Yulun Zhang
Generative Models Efficient ML Computer Vision
  • WorldDynCache addresses the inefficiencies of existing caching methods in diffusion world models.
  • The framework includes a risk estimator to track and calibrate approximation defects.
  • It employs a condition- and phase-aware surrogate for latent evolution approximation.
  • WorldDynCache achieves significant speedups in inference time while preserving high generation quality.
Read more
Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers
Farbod Faraji, Francesco Belardinelli
Interpretability
  • Introduction of a verifier-guided workflow for symbolic model discovery.
  • VG outperforms traditional methods in generalization across initial conditions.
  • Successful transfer to high-dimensional physical data, particularly in vortex shedding scenarios.
  • Demonstration of the importance of latent dynamics compatibility with pretraining.
Read more
CoSynFlow: Conformal Symplectic Neural Flows for Cross-System Prediction of Dissipative Hamiltonian Dynamics
Baige Xu, Takaharu Yaguchi
Theory Optimization Time Series
  • CoSynFlow maintains exact conformal symplecticity for dissipative Hamiltonian dynamics.
  • The model enables continuous-time flow learning, allowing predictions for any time t in a single evaluation.
  • It facilitates cross-system prediction, allowing a single model to generalize to unseen systems without retraining.
  • The architecture preserves the geometric structure of the dynamics by design, addressing limitations of existing neural operator methods.
Read more
Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language
David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Ranković, Philippe Schwaller
NLP Multimodal
  • CheMatE is a bi-semantic embedding model that captures both SMILES and natural language in a unified representation space.
  • The model employs a two-stage training process, combining masked language modeling with contrastive learning.
  • CheMatE is trained on a large corpus of SMILES-annotated scientific documents, enhancing its contextual understanding.
  • The model achieves strong performance across molecular property prediction and scientific language understanding tasks.
Read more
On the Implicit Flatness Bias of Sharpness-Aware Minimization: A Linear Stability Analysis with Quantitative Hyperparameter Bounds
Jiaxin Deng, Junbiao Pang
Optimization Theory
  • Introduces a linear stability framework for analyzing SAM's implicit flatness bias.
  • Establishes a quantitative relationship between perturbation radius, batch size, learning rate, and the sharpness of minima.
  • Demonstrates the trade-off between perturbation radius and stable training.
  • Validates theoretical predictions through extensive experiments on CIFAR-100.
Read more
Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms
Tezan Sahu, Aritra Das, Pankaj Mittal, Sudipta Das
NLP Large Language Models Efficient ML
  • Introduces a capability-taxonomy-driven pipeline for curating regression evaluation sets.
  • Addresses the limitations of existing customer-side evaluation frameworks.
  • Utilizes a hybrid classifier and Invocation Quality rater to assess query capabilities.
  • Implements a rule-based consolidator for decision-making on query inclusion.
Read more
CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning
Jian Zhang, Bingyi Wang, Yizhi Liu
NLP Large Language Models Reinforcement Learning
  • CausalOPD introduces a curriculum online process distillation framework for causal reasoning.
  • The method focuses on identifying the first wrong step in reasoning chains to correct errors effectively.
  • Improvements in average path correctness by 23.4 percentage points over standard methods.
  • Significant reduction in right-label-wrong-reasoning (RLWR) rate from 15.7% to 4.4%.
Read more
Latent-Regime Bias Auditing for Volatility Forecasting
Arthur Chagas, Pedro Bento, Yan Aquino, Arthur Buzelin, Wagner Meira Jr., Cristiano Arbex Valle
Time Series
  • Volatility forecasting should be evaluated as a conditional reliability problem rather than solely through aggregate accuracy metrics.
  • The proposed audit framework identifies hidden regime biases and tail underprediction in volatility forecasts.
  • Empirical results show that models with competitive aggregate accuracy can still fail in specific market regimes.
  • The framework allows for a nuanced understanding of forecast reliability across different economic conditions.
Read more
Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views
Jiaqi Xiong, Yuntao Hu, Yu Zheng, Yifei Shi, Xinyue Guo, Jiaxin Qi
Graph Learning
  • Introduces CoCoS, a contrastive pretraining framework for single-cell transcriptomics.
  • Addresses the gap in existing models that focus on gene-level reconstruction rather than whole-cell representation.
  • Implements co-expression-guided gene partitioning and expression-aware contrast-set construction to enhance model training.
  • Demonstrates superior performance in downstream tasks such as cell-type annotation and gene regulatory network inference.
Read more
Trajectory inference via Acceleration Matching
Bartolo Dazzini, Giovanni Conforti, Alain Durmus, Aram-Alexandre Pooladian
Time Series Generative Models Theory
  • Introduction of Acceleration Matching (AM) as a novel algorithm for trajectory inference.
  • AM avoids the computational burdens of existing methods by not requiring simulation during training.
  • The algorithm generates smooth trajectories by targeting an explicit acceleration field.
  • Numerical experiments show AM's competitive performance against existing algorithms.
Read more
Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models
Lele Zheng, Weifeng Kong, Xinyi Zhang, Ke Cheng, Tao Zhang, Yulong Shen
NLP Large Language Models Optimization
  • Introduces SAGE, a noise-aware shrinkage method for adaptive update scaling in DP-ZO.
  • Demonstrates that fixed-scale updates do not account for the non-stationary reliability of privatized gradient estimates.
  • Theoretical analysis shows SAGE effectively reduces update-risk while preserving useful descent.
  • Empirical results indicate SAGE outperforms existing methods under identical privacy constraints.
Read more
Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard
Hector Zenil, Luan Ozelim
NLP Large Language Models Theory
  • Introduction of F-ICL benchmark for evaluating algorithmic reasoning in LLMs.
  • Exhaustive enumeration of programs allows for the computation of a Bayes-optimal standard.
  • Most evaluated models perform worse than a simple keystroke reference despite high accuracy.
  • The study reveals a significant gap between model performance and the Bayes-optimal standard.
Read more
Sharp Root Anti-Concentration via Projective Incidence and Ordered Root Laws
Zijun Wang, Yuchen Miao, Yifan Hu, Huanmin Liu
Theory Optimization Graph Learning
  • Establishes a dimension-free characterization of interval-hitting constants for piecewise-Lipschitz functions.
  • Removes the previous √N loss in the context of cube-supported coefficients.
  • Provides necessary and sufficient conditions for finite interval-hitting constants for monic polynomials.
  • Introduces coefficient-based tests for dependent and singular laws.
Read more
Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces
Francesca Carlon, Vincent Ginis, Andres Algaba
NLP Large Language Models Efficient ML
  • Shortening reasoning traces can reduce costs and latency in LLMs.
  • A numeric/concision prompt effectively shortens reasoning without consistent accuracy gains.
  • Concise instructions can improve accuracy, particularly under token constraints.
  • Lower-effort reasoning often results in higher accuracy compared to high-effort reasoning.
Read more
Tight Worst-Case Bounds for the Smallest Eigenvalue of ReLU NTK Gram Matrices
Zhao Song
Theory Optimization
  • Establishes a dimension-free lower bound for the smallest eigenvalue of ReLU NTK Gram matrices.
  • Proves that Ξ»min(H) is proportional to projective separation βˆ†Β± divided by √log n.
  • Constructs worst-case examples that match the upper bound, demonstrating tightness of the results.
  • Highlights the importance of geometric nondegeneracy in neural network optimization.
Read more
Neural operator learning for collision-aware trajectory planning of spacecraft swarms
Sidhdharth D. Sikka, Suyi Gao, Zehui Lu, Rongjie Lai, Shaoshuai Mou
Robotics Optimization
  • Introduces a permutation-equivariant neural operator for collision-aware trajectory planning.
  • Trained without optimal trajectory labels using self-supervised physics objectives.
  • Achieves zero-shot generalization to swarms of 1,000 spacecraft in dense debris fields.
  • Matches the accuracy of traditional optimal-control methods while improving swarm proximity.
Read more
Adaptive Quantum Physics-Informed Neural Networks for Differential Equations with Applications to Fluid Dynamics
Fabio Pereira dos Santos, Renato Portugal, JΓΊlio de Castro Vargas Fernandes, Lucas Timotheo Sanches
Theory Optimization Efficient ML
  • Introduction of adaptive collocation point sampling to improve accuracy in QPINNs.
  • Identification of optimization as a critical bottleneck in QPINNs, alongside expressivity.
  • Development of a trainable loss-weighting scheme for balanced training.
  • Demonstration of at least 60% improvement in solution accuracy for specific fluid dynamics problems.
Read more
To Describe or Construct Statistical Learning Models Using the Category-theoretical Language
Congwei Song
Theory
  • Utilizes category theory to describe statistical learning models.
  • Introduces a unified descriptive approach for constructing complex models.
  • Highlights the construction of the Transformer model as a key example.
  • Aims to engage researchers from diverse fields in statistical learning research.
Read more
Scaling an Autoregressive Transformer for Single-Cell Generation
Aleksandr Sharipov, Yusif Mukhtarov, Igor Molybog
Generative Models
  • Introduction of a self-supervised autoregressive transformer for generating single-cell gene expression vectors.
  • Identification of a joint two-exponent scaling law for model size and training data size in single-cell models.
  • Demonstration of the correlation between pretraining loss and downstream evaluation metrics.
  • Potential applications for fine-tuning the model for perturbation response prediction.
Read more
Fused Bayesian Flow Networks for Dual-Target Molecular Design
Jingyuan Zhou, Shikui Tu, Lei Xu
Generative Models
  • Introduction of FusedBFN, a fused Bayesian flow network for dual-target molecular design.
  • Utilizes a product-of-experts formulation to integrate information from two target proteins.
  • Leverages a pretrained target-aware BFN to address the scarcity of dual-target structural data.
  • Implements a chemically aware alignment strategy for better integration of structural features.
Read more
Sphere Retraction Normalizations
Jie Zhang, Cheng-Fang Su, Yi-Jui Huang, Min-Te Sun
Theory Optimization Large Language Models
  • Unification of residual connections and GeoNorm as retraction-based updates.
  • Introduction of two new retraction methods: Proj-SpheretNorm and Cay-SpheretNorm.
  • Development of a generalized p-SpheretNorm that encompasses both new methods and GeoNorm.
  • Empirical results show Proj-SpheretNorm achieves superior performance on nanoGPT models.
Read more
ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference
Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri
Large Language Models Efficient ML
  • Introduction of a per-RoPE-wavelength distance window for efficient LLM inference.
  • Achieves 37-48% pruning of query-key inner-product terms with high accuracy retention.
  • Maintains performance across various long-context benchmarks.
  • Implemented with minimal modifications to existing FlashAttention kernels.
Read more
AlphaG-OPD: Reliability-Gated Sibling Counterfactuals for On-Policy Distillation in Symbolic Alpha Factor Discovery
Yaoyu Su
Reinforcement Learning Generative Models Optimization
  • Introduces AlphaG-OPD, a structural on-policy distillation framework for symbolic alpha factor discovery.
  • Separates decision-making into three components: where to teach, what to teach, and how to consolidate teaching.
  • Employs reliability gating to ensure only reliable sibling actions are taught, enhancing learning efficiency.
  • Maintains the global objective of Trajectory Balance while providing local action guidance.
Read more
Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents
Yidan Lin, Kaixiang Wang, Jiong Lou, Jie Li
Large Language Models Efficient ML NLP
  • Router-Mem introduces a novel framework for long-horizon agent memory that balances efficiency and answer quality.
  • The sufficiency router enables early termination of processing when sufficient evidence is available, reducing unnecessary computation.
  • The framework shows strong performance on benchmark datasets, achieving high scores while decreasing average inference time.
  • Router-Mem effectively combines retrieval and deeper analysis, optimizing resource allocation based on query needs.
Read more
GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits
Nitish Nagesh, Elahe Khatibi, Thomas Dean Hughes, Mahdi Bagheri, Pratik Gajane, Amir M. Rahmani
Graph Learning Large Language Models Interpretability
  • GoT-CD generates multiple candidate graphs in parallel, enhancing the robustness of causal discovery.
  • The method ensures that critical pathways for fairness audits are preserved, addressing a significant gap in existing methodologies.
  • GoT-CD outperforms traditional causal discovery methods and LLM-based approaches in terms of structural validity and path-specific fairness.
  • The study emphasizes the need for causal discovery methods to be evaluated based on their ability to recover relevant pathways for fairness analysis.
Read more
Federated generative event models for tokenized electronic health records
Michael C. Burkhart, Luke Solo, Inhyeok Lee, S'Khaja Charles, Zewei "Whiskey" Liao, Kaveri Chhikara, Dema Therese, Wan-Ting Liao, Catherine A. Gao, William F. Parker, Brett K. Beaulieu-Jones
Federated Learning Generative Models Time Series
  • GEM representations are significantly more portable across health systems compared to conventional supervised models.
  • Federated learning approaches recover most of the performance of centralized training, especially for data-limited sites.
  • The primary challenge is cross-site transfer and adaptation rather than the federated optimization process.
  • Centralized multi-site training offers limited improvements over strong local models when large datasets are available.
Read more
Contrast-invariant deep ptychography neural networks
Albert Vong, Steven Henke, Oliver Hoidn, Hanna Ruth, Junjing Deng, Apurva Mehta, David Shapiro, Alexander Hexemer, Nicholas Schwarz
Computer Vision Optimization Efficient ML
  • Introduces a factorization strategy to decouple object texture from measurement scaling.
  • Utilizes a real/imaginary output representation for improved optimization stability.
  • Implements a dynamic scaling factorization optimized during test time.
  • Demonstrates a 5x reduction in Fourier error across various experimental datasets.
Read more
Quantization Effects on Biomedical LLM Reliability
Anton Rasmussen, Hong Qin
NLP Large Language Models
  • Scoring rule sensitivity can reverse calibration rankings without affecting accuracy.
  • Prompt template choice can lead to significant accuracy variations, comparable to model choice.
  • INT8 quantization has minimal impact on specialized models, while INT4 effects are heterogeneous.
  • Post-hoc temperature scaling can improve calibration but is specific to certain scoring rules.
Read more
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Ranxu Zhang, Guinan Chen, Chenshaodong, Jinghao Lin, Xiaozhou Xu, Sunzhe, Yanyong Zhang, Chao Wang
Reinforcement Learning Large Language Models NLP
  • ADRS provides a solution to the sparse feedback problem in agentic reinforcement learning by enabling token-level credit assignment.
  • The framework incorporates teacher confidence and realized returns to modulate reward shaping effectively.
  • Experiments show consistent performance improvements across various interactive benchmarks and RL backbones.
  • ADRS maintains skill-free rollouts during inference, enhancing its applicability in real-world scenarios.
Read more
Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift
Hanyu Su, Carlota Julbe i Juanola, Yibo Hu
Computer Vision
  • Calibration is crucial for evaluating subtype robustness, not just accuracy.
  • Models exhibit overconfidence in predictions for unseen subtypes despite accuracy loss.
  • Generic image corruption leads to a more significant drop in confidence compared to subtype novelty.
  • Recalibration methods do not fully recover lost calibration under unseen subtype conditions.
Read more
ConformalShift: Targeted Event Reordering Against Adaptive ECG Monitoring
Arash Vashagh, Yasmin Vashagh
Time Series
  • ConformalShift is an event-reordering attack that targets adaptive ECG monitoring systems.
  • The attack suppresses the detection of ventricular ectopic beats without modifying ECG data or classifier outputs.
  • Experimental results show a suppression rate of 66.7% for Extra Trees and 60.0% for HistGradientBoosting classifiers.
  • The study highlights the importance of event order in adaptive conformal prediction systems.
Read more
GENESIS: Towards Explainable Causal Discovery
Abhinav Thorat, Ravi Kumar Kolla, Vishak K Bhat, Harsh Vardhan Singh Chauhan, Niranjan Pedanekar
Graph Learning Interpretability Large Language Models
  • Genesis ensures decision traceability for each edge in the causal graph.
  • The framework integrates motif-level reasoning with statistical validation.
  • Achieves 100% decision traceability across various experimental settings.
  • Outperforms traditional statistical methods in terms of Structural Hamming Distance.
Read more
ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection
Camile Lendering, Erkut Akdag, JoaquΓ­n Figueira, Egor Bondarev
Computer Vision Generative Models
  • Introduces ReFP-AD, a method for stable anomaly detection in high-dimensional token spaces.
  • Addresses instability in Energy-Based Models due to geometric factors like anisotropy.
  • Achieves 98.6% Image AUROC on MVTec-AD and 97.3% on VisA, outperforming prior methods.
  • Demonstrates the critical role of geometric reparameterization for effective MCMC sampling.
Read more
Maglev: Sliding Recurrent Memory
Bo Liu, Qiang Liu
NLP Large Language Models Efficient ML
  • Maglev combines sliding-window attention with a recurrent memory mechanism for improved language modeling.
  • The architecture allows for parallel training through a prefiller that generates memory targets independently of the decoder.
  • A memory consistency loss is employed to align the decoder's memory with the prefiller's outputs.
  • Empirical results show significant improvements in validation loss and accuracy on downstream tasks.
Read more
ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces
Kunal Kumar Pant, Nithin Nagaraj
NLP Large Language Models Theory
  • ChaosProbe offers a new method for analyzing frozen transformer input-embedding spaces using neurochaotic principles.
  • The method generates fixed-length signatures that summarize the response characteristics of embedding matrices.
  • The study validates the effectiveness of ChaosProbe through correlation analyses among different transformer models.
  • Results indicate that the embedding spaces contain meaningful geometric structures that can be analyzed independently of contextual computations.
Read more
FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs
Amin Farajzadeh, Melike Erol-Kantarci
Reinforcement Learning Federated Learning Efficient ML
  • FedCritic-MIMO enables communication-efficient coordination among independently deployed RAN controllers.
  • The framework allows local execution of actors while facilitating shared critic parameter exchange.
  • Achieves significant performance improvements in network throughput and quality-of-service metrics.
  • Reduces critic communication overhead by about 76% compared to traditional methods.
Read more
LLMs Can Annotate Attribution Graphs
Ameen Patel, Max Zhang, Nathan Hu
Large Language Models Interpretability
  • Introduction of a low-cost pipeline for annotating attribution graphs using LLMs.
  • Validation of the interpretability of LLM-generated supernodes against human annotations.
  • Successful application of the pipeline on a two-hop Capitals task with high recovery rates.
  • Proof of concept for open-ended exploration of attribution graphs through automated annotation.
Read more
HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning
Haowei Liu, Jiamian Wang, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang
Reinforcement Learning Large Language Models NLP
  • Introduces trajectory-level hindsight critique (TLHC) to enhance search-augmented RL training.
  • Achieves 39.4% average EM on a standard benchmark suite, outperforming prior methods.
  • Demonstrates the critical role of hindsight in improving agent performance.
  • Utilizes a frozen judge to provide directive critiques of failed trajectories.
Read more
CoRe-GNN: Multilevel Message passing on Coarsened graphs
Antonin Joly, Nicolas Keriven, Aline Roumy
Graph Learning
  • CoRe-GNN combines inter-cluster and intra-cluster message passing to leverage the strengths of both approaches.
  • The architecture is agnostic to the choice of coarsening algorithm, allowing for flexibility in implementation.
  • A novel batching scheme enables CoRe-GNN to efficiently handle large graphs with millions of nodes.
  • The method shows competitive performance on various node classification tasks, outperforming existing baselines.
Read more
FL-OA: A Byzantine-Robust Federated Learning Framework with Outsourced Auditing for Intelligent Devices
Hongliang Zhang, Zhongyuan Yu, Fenghua Xu, Teng Hu, Jian Meng, Jiguo Yu
Federated Learning
  • FL-OA utilizes outsourced auditing to enhance Byzantine robustness without strong assumptions.
  • The framework mitigates divergence among benign updates through innovative local training techniques.
  • A parameter importance indicator is introduced to address the curse of dimensionality.
  • Theoretical analysis supports the framework's design and effectiveness.
Read more
UNVaMP: Neural Knowledge Tracing with Variational Regularization of Latent Knowledge Dynamics
Carson J. Cook, Ahmed J. Zerouali, Anthony Schmidt, Reginald Ziedzor, Paul Lin, Luke G. Eglington
Theory Interpretability Time Series
  • UNVaMP integrates observed interactions with internal memory for dynamic knowledge representation.
  • The architecture supports both purely neural and hybrid configurations for flexibility and interpretability.
  • UNVaMP-MLP shows superior predictive performance compared to other models on multiple datasets.
  • The model allows for controlled volatility in knowledge estimates and quantifies uncertainty.
Read more
Approximate Speculative Decoding
Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang
NLP Large Language Models Efficient ML
  • ASD replaces binary first-mismatch truncation with budgeted longest-prefix selection, allowing for more flexible acceptance of mismatches.
  • The method does not require a new draft model or fine-tuning, making it easy to implement.
  • ASD improves throughput by 3.05% to 15.26% compared to strict verification methods across multiple tasks.
  • The approach allows for the reuse of target-greedy suffixes, enhancing efficiency without additional computational costs.
Read more
Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality
Shengzhi Deng, Chenqi Ye, Yanze Guo
Generative Models Optimization Theory
  • Pseudorandom streams can be treated as learnable inputs in diffusion models.
  • The study distinguishes between next-value predictability and exploitability within the target diffusion system.
  • Different pseudorandom orbits lead to significant variations in generation quality and diffusion losses.
  • The relationship between pseudorandom inputs and learning systems follows an empirical power law.
Read more
Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving
Xiang Li, Pengcheng Wang, Huazheng Wang, Saurabh Bagchi
Large Language Models Efficient ML
  • SALT introduces a matrix cosine regularizer to align task subspaces into a unified basis.
  • The framework decouples representational capacity from physical memory costs, allowing for efficient multi-tenant serving.
  • SALT achieves up to 18.5% absolute accuracy gains over existing compression methods while reducing memory usage by up to 16x.
  • The method improves serving throughput by up to 51% under PCIe bandwidth pressure and 28% under GPU VRAM constraints.
Read more
PatTree: a novel approach for automated creation of multimodal, graph-based patient representations for medical classification tasks
Julia Gehrmann, Lars Quakulinski, Hamza Naseem, Oya Beyan
Graph Learning Multimodal
  • PatTree automates the integration of multimodal, longitudinal patient data.
  • The approach creates a tree-based patient representation without requiring data harmonization.
  • PatTree preserves semantic relationships in clinical data, enhancing interoperability.
  • Achieved state-of-the-art classification performance in distinguishing cognitive states.
Read more
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, Mohammad Mahdi Kamani, Ritika Borkar, Makesh Tarun Chandran, Pantea Zardoshti, Bita Darvish Rouhani
Large Language Models Efficient ML Optimization
  • Introduction of cross-model KV cache transfer to reduce prefill costs in LLMs.
  • Discovery of substantial linear structure in KV relationships across model families.
  • Development of a closed-form ridge mapper that operates efficiently on a per-head basis.
  • Validation of the method across multiple model families, achieving high accuracy retention.
Read more
A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation
Abdallah Khemais
Theory Interpretability
  • Establishes conditions for deterministic output collapse when carriers are deleted from a neural network.
  • Demonstrates that activation patching and weight-space ablation produce different effects on network behavior.
  • Derives an exact interaction term for a specific architecture, revealing the limitations of treating components as independent.
  • Validates theoretical predictions with synthetic experiments on small transformers, showing strong rank correlation with model behavior.
Read more
Sedentary Behavior Classification for Wearable Sensors with a CNN-BiLSTM Model
Yuliang Chen, Weiwei Shi, Jingjing Zou, Rong Zablocki, Animesh Kumar, Jordan A. Carlson, Sheri J. Hartman, Mikael Anne Greenwood-Hickman, Paul R. Hibbing, Marta Jankowska, Jay Yang, Arun Kumar, Loki Natarajan
Time Series
  • The CHAP model effectively classifies sedentary behavior using hip-worn accelerometer data.
  • Transferability of the hip-trained model to wrist-worn data shows a decrease in accuracy, necessitating adaptation.
  • Fine-tuning the model with wrist data improves performance compared to training from scratch.
  • The study emphasizes the importance of posture in accurately measuring sedentary behavior.
Read more
Deep Learning-Based Estimation of Ground Reaction Forces in Parkinsonian Gait Using an Optimized Set of IMU Data
Run Lin, Yingtian Tang, Jiawen Xu, Dongfei Huo, Lefan Wang, Helen Dawes, Dominic J. Farris, Dong Wang, Xijin Hua
Time Series
  • Introduces a deep learning framework for estimating ground reaction forces in Parkinsonian gait.
  • Achieves high accuracy in estimating vGRFs using a minimal configuration of wearable IMUs.
  • Demonstrates significant differences in optimal sensor placement between PD patients and healthy controls.
  • Identifies a compact setup that allows for robust estimation with only two IMUs.
Read more
NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory
Yizhe Chen, Wenshuai Yao, Saiya Wang, Yuannuo Feng, Wenbo Qi, Kechao Tang, Ngai Wong, Wenyong Zhou, Wang Kang
Efficient ML
  • NANQ optimizes quantization by considering hardware noise, improving precision allocation.
  • The framework assigns finer quantization resolution in low-noise areas and avoids redundancy in high-noise regions.
  • Layer-wise bit-widths are determined based on each layer's saturation point under noise, enhancing model performance.
  • Experimental results show significant accuracy improvements in vision and language models compared to existing quantization methods.
Read more
Sparse Weight Decomposition for Efficient Circuit Extraction
Chuanhao Yan, Xuhan Huang, Yawen Duan, Zhenfei Yin, Hang Zhao, Bryan Dai, Jie Fu
Interpretability Large Language Models Efficient ML
  • SWD allows for efficient circuit extraction from pretrained models without additional training.
  • The method matches or exceeds the fidelity of existing approaches while using less than 1% of the data.
  • SWD supports full-model replacement of attention and MLP weight matrices after fine-tuning.
  • A zero-data variant of SWD enables broader use in mechanistic interpretability analysis.
Read more
DiagLoop: A Counterfactual Data Flywheel with Stage-Localized Reinforcement for Diagnostic LLMs
Jian Zhang, Bingyi Wang, Yizhi Liu
Large Language Models Reinforcement Learning NLP
  • DiagLoop synthesizes counterfactual scenarios to enhance training data for diagnostic LLMs.
  • The model employs a unique stage-localized reinforcement learning approach to improve reasoning accuracy.
  • It achieves significant improvements in path correctness over conventional models in various diagnostic contexts.
  • The methodology allows for local deployment, addressing privacy and latency concerns in sensitive data environments.
Read more
Population-Robust Feature Selection via Generalized Welfare Optimization
Ruiqi Lyu, Alistair Turcan, Bryan Wilder
Optimization Efficient ML Interpretability
  • Introduction of PopFS for robust feature selection across diverse populations.
  • Utilization of a tunable welfare objective to balance predictive performance and protection for lower-benefit populations.
  • Scalable optimization strategy using multitask sparse learning and ranked refit search.
  • Demonstrated strong performance improvements in both average and worst-case scenarios.
Read more
Adaptive Sampling for Automated Post-Disaster Rapid Damage Assessment via Level-Set Cost-Aware Bayesian Optimization
Boyang Xu, Mostafa Reisi Gahrooei, Mohammad Ilbeigi, Hao Yan
Optimization Robotics Efficient ML
  • Introduction of a novel Ordinal Deep Kernel Gaussian Process (ODGP) for modeling spatial correlations in damage assessment.
  • Development of a cost-aware level-set acquisition function that optimizes data collection locations based on information gain and travel costs.
  • Validation of the framework through synthetic studies and real disaster data, showcasing its effectiveness in rapid damage assessment.
  • Demonstration of reduced uncertainty in damage predictions and improved operational efficiency for emergency response.
Read more
SAKI: Score-Aware Low-Rank Key Indexing for Long-Context KV Retrieval
Lin Zhang
NLP Large Language Models Efficient ML
  • SAKI optimizes key compression specifically for attention score fidelity, improving retrieval performance.
  • The method outperforms key-PCA in all tested models, reducing recall error by 13-30%.
  • The theoretical framework predicts empirical results with high accuracy, validating the approach.
  • SAKI is training-free and provides a significant improvement in deep layers of transformer models.
Read more
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
Xujun Che, Yuchen Yuan, Weida Zhao, Chenyang Yu
Reinforcement Learning Large Language Models Theory
  • Abstention as a discrete action can lead to the collapse of both reward gradients and KL anchors in reinforcement learning.
  • The collapse occurs under specific conditions, including bounded readouts and the penalty for incorrect answers.
  • A structural repair method is proposed, moving abstention to a confidence reporting mechanism that avoids shared saturation factors.
  • Simulations and experiments confirm the predictions, demonstrating improved model performance through the proposed method.
Read more
Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments
Muhammad Faizan Raza, Shuo (Luna) Yang, Satish Mahadevan Srinivasan, Joanna F. DeFranco
Large Language Models NLP Generative Models
  • Introduction of a unified LLM operations (LLMOps) architecture for real-time deployments.
  • Integration of real-time data ingestion and continual learning to address knowledge staleness and hallucinations.
  • Framework designed to meet regulatory compliance and operational demands in high-risk sectors.
  • Utilization of software design patterns to optimize latency, cost, and accuracy trade-offs.
Read more
SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation
Zihuan Qiu, Zhiyang Liao, Chiyuan He, Yi Xu, Fanman Meng, Linfeng Xu, Qingbo Wu, Hongliang Li
Computer Vision NLP Efficient ML
  • SAFE-Merge introduces a risk-aware sparse masking approach to select safe parameter updates for merging.
  • The framework employs masked low-rank recovery to restore task-specific information without altering masked parameters.
  • It achieves the best H-score across multiple vision and language benchmarks, demonstrating effective continual merging.
  • The method incurs no additional inference costs, making it efficient for real-world applications.
Read more
Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning
Botao Dong, Longyang Huang, Ning Pang, Hongtian Chen
Reinforcement Learning Generative Models Robotics
  • Introduces BAC-PE to correct Q-value estimations in offline RL.
  • Theoretically analyzes the convergence of BAC-PE and provides an upper bound on Q-function differences.
  • Employs diffusion models for effective policy regularization and distribution matching.
  • DPBAC algorithm shows superior performance on D4RL tasks compared to existing methods.
Read more