AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

48 Papers today
8h Update frequency
7 Days of history
Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling
Pedro Sousa, Will Tebbutt, Sadiq Jaffer, Robin Young, Anil Madhavapeddy, Richard E. Turner
Time Series
  • Earth observation embeddings can replace hand-crafted topographic descriptors in weather downscaling.
  • The proposed method improves probabilistic skill for temperature and wind speed predictions.
  • The approach is effective across diverse climatic regions and for previously unobserved locations.
  • Improvements in CRPS skill demonstrate the potential of using long-timescale embeddings for short-timescale predictions.
Read more
FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data
Viktoria Schuster, Sana Tonekaboni, Caroline Uhler
Multimodal Optimization Interpretability
  • FiGuRO provides a dynamic approach to intrinsic dimension estimation in multi-modal data.
  • The framework allows for disentanglement of shared and private information as an emergent property.
  • FiGuRO outperforms existing ID estimation methods and is robust to hyperparameter changes.
  • The method captures distinct ID scales and varying subspace ratios effectively.
Read more
Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing
Ziqiang Li, Yun Liu, Gouhei Tanaka
Time Series Efficient ML
  • DTW-GBC organizes training samples into granular balls for efficient classification.
  • The method reduces the influence of mislabeled samples on classification performance.
  • Two granular-ball construction strategies are proposed: random splitting and label-informed splitting.
  • Experiments show DTW-GBC outperforms traditional DTW-based classifiers in noisy label scenarios.
Read more
Diffract: Spectral View of LLM Domain Adaptation
Nikita Borodin, Maria Krylova, Artem Zabolotnyi, Dmitry Aspisov, Egor Shikov, Nikita Tyuplyaev, Oleg Travkin, Roman Alferov, Dmitry Vinichenko
NLP Large Language Models
  • CPT maintains singular value spectra while adaptation is driven by singular vector changes.
  • Significant domain-dependent heterogeneity in attention heads allows for selective updates.
  • Up to 60% of head updates can be removed without quality loss, improving accuracy by up to 4%.
  • Linear interpolation between CPT checkpoints shows smooth domain-quality transitions.
Read more
CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster Sampling
Alexander SchiΓΈtz, Bertram Hage, Christian Rand, Felix Thomsen, Peder Heiselberg
Time Series
  • CRHT addresses geographic bias and navigational realism in vessel trajectory prediction.
  • An online K-means cluster sampling strategy is introduced to enhance training diversity.
  • The hybrid architecture integrates local kinematic feature extraction with global attention.
  • CRHT achieves superior performance in short-term forecasting compared to existing models.
Read more
Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models
Myeong-Ju Cho, Hye-Bin Shin, Seo-Hyun Lee, Seong-Whan Lee
Multimodal Time Series NLP
  • BLPM aligns continuous EEG representations with language semantics through semantic embedding prediction.
  • The CELP encoder promotes higher-level abstraction by predicting latent representations from contextual EEG observations.
  • The MQSD module enables selective access to distinct semantic factors within EEG segments based on task relevance.
  • BLPM avoids the pitfalls of discrete tokenization and autoregressive generation, enhancing EEG decoding capabilities.
Read more
Task- and dataset-specific information in protein language models
Roman Joeres, Ilya Senatorov, Olga V. Kalinina
NLP
  • Intermediate layers of PLMs often provide more informative embeddings than the last layer for specific downstream tasks.
  • The relationship between the pre-training objective and downstream task influences the distribution of relevant information across PLM layers.
  • Dataset characteristics play a crucial role in determining the effectiveness of PLM embeddings for whole-protein tasks.
  • Performance of PLMs significantly drops when applied to artificial protein sequences, underscoring the importance of authentic training data.
Read more
Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment
Yan Wang, Chuan-Xian Ren
Graph Learning Optimization Theory
  • Introduces a pair-centric approach to graph rewiring focused on communication shortages.
  • Develops a shortage score that ranks node pairs based on their structural demand versus current support.
  • Utilizes Optimal Transport to optimize edge additions and budget allocation for effective communication.
  • Demonstrates significant improvements in performance on standard graph benchmarks.
Read more
Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
Zijian Zhao, Sen Li
Theory Multimodal Interpretability
  • Introduces a unified theory for multiplicative dual-encoder networks across various domains.
  • Defines low interaction rank functions and their interaction spectrum to measure complexity.
  • Addresses approximation error, identifiability issues, and sample complexity in encoder networks.
  • Demonstrates that normalization techniques can enhance interpretability of learned representations.
Read more
Confidence Calibration of Deep Learning Systems
Coby Penso
Theory Efficient ML Interpretability
  • Introduces novel methods for confidence calibration in deep learning systems under noisy labels.
  • Explores conformal prediction techniques that are robust to label noise.
  • Presents local differential privacy approaches for conformal prediction.
  • Demonstrates effective unsupervised target domain calibration methods.
Read more
Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction
Yi Liu
Theory Optimization Robotics
  • Introduces a decision-resource framework that utilizes terminal symmetry for construction tasks.
  • Develops SYMBUILD, which implements a statewise refinement approach to improve decision-making.
  • Demonstrates significant improvements in verified efficiency across multiple domains.
  • Achieves the lowest mean capped verifier cost compared to other planners in GRN OOD scenarios.
Read more
Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review
Christopher Braun, Julian Raible, Marco F. Huber
Theory
  • PIML integrates physical knowledge into ML to enhance reliability and interpretability in PHM.
  • The review categorizes existing studies into four classes based on biases and approaches.
  • PIML shows improved predictive performance over traditional methods, particularly in specific applications.
  • There is a need for more evidence supporting claims of enhanced interpretability and generalization.
Read more
LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
Zhixin Zhang, Xinke Jiang, Zhibang Yang, Weixuan Xu, Guohong Qiu, Xu Chu, Junfeng Zhao, Yasha Wang
NLP Large Language Models Reinforcement Learning
  • LoongReflect formulates reflection as a memory-control policy to enhance long-horizon reasoning.
  • The framework utilizes a reversible trajectory tree with explicit actions for reflection and backtracking.
  • A dual-channel learning approach combines privileged teacher feedback with outcome-based reinforcement learning.
  • Experiments show significant improvements over traditional reinforcement learning methods.
Read more
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
Karim Aly, Alexei Sharpanskykh, Jacco Hoekstra
Generative Models
  • TailBooster effectively addresses the rarity of extreme events in historical flight data.
  • The framework combines generative modeling with anomaly detection to ensure operational validity.
  • Significant improvements in prediction accuracy for extreme events were achieved compared to conventional methods.
  • TailBooster is adaptable to various domains beyond aviation, where extreme-event prediction is necessary.
Read more
A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes
Yijin Ni, Xiaoming Huo
Theory Efficient ML
  • Introduces a joint-distribution approach to fair representation learning for continuous sensitive attributes.
  • Proposes a new discrepancy measure that avoids the need for conditional laws, enhancing computational efficiency.
  • Demonstrates that the HSIC estimator converges faster than traditional nonparametric estimators.
  • Establishes theoretical connections between the joint discrepancy and existing fairness criteria.
Read more
Clustered Randomized Smoothing for Stochastic Prediction Functions
Eduardo Figueiredo, Frederik Mathiesen, Julian Schumann, Jens Kober, Arkady Zgonnikov, Luca Laurenti
Robotics Generative Models Reinforcement Learning
  • Introduction of clustered Ξ±-smoothing to enhance robustness in stochastic predictors.
  • Local Ξ±-smoothing within clusters prevents mode collapse in multi-modal distributions.
  • The framework is flexible and agnostic to the clustering algorithm used.
  • Empirical results show a 27% reduction in Wasserstein distance and an 81% reduction in collision rates compared to state-of-the-art methods.
Read more
Mapping and Measuring the Behavioral Evolution of Large Language Models
Dong Qiao, Chris Ding, Jicong Fan
NLP Large Language Models
  • Introduces three novel sentence-level dissimilarity measures for analyzing LLM behavior.
  • Demonstrates coherent clustering of model families and decreasing cross-family distances over time.
  • Validates findings with a token-level analysis, confirming the robustness of the results.
  • Establishes a mathematical relationship between behavioral similarity and training dynamics.
Read more
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning
Chenhua Fan, Jiahui Zhu, Yuhang Zhang, Honghao Wei
Reinforcement Learning Optimization Theory
  • BSPG explicitly separates reward improvement and boundary regulation in policy updates.
  • The method achieves a finite-horizon O(1/√T) convergence bound for constraint residuals.
  • BSPG demonstrates superior performance in reward maximization and boundary tracking compared to baseline methods.
Read more
SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks
Yue Xia, Tayyebeh Jahani-Nezhad, Mayank Bakshi, Rawad Bitar
Federated Learning Efficient ML NLP
  • SeFoRA addresses the bilinear mismatch problem in federated LoRA by using sketch aggregation.
  • The algorithm allows for heterogeneous client ranks, enabling clients with different computational budgets to participate effectively.
  • SeFoRA-Ho provides a rank-homogeneous solution for direct adapter aggregation.
  • Convergence to a first-order stationary point is proven for the rank-homogeneous setting.
Read more
Let it Cook: Learning to Wait in Sequential Decision Making
Christopher Watson, Arjun Krishna, Dinesh Jayaraman, Rajeev Alur
Reinforcement Learning Robotics Optimization
  • Introduces a waiting policy for sequential decision making to optimize resource usage.
  • Formulates 'learning to wait' as a multi-objective optimization problem.
  • Employs a lexicographic MORL algorithm to train agents across various environments.
  • Demonstrates significant waiting periods (over 50% of task duration) without performance loss.
Read more
Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique
Sanidhya Vijayvargiya, Rahul Lokesh
NLP Large Language Models Efficient ML
  • Introduction of the Latent Critic, a low-latency mechanism for hallucination detection in LLMs.
  • Demonstrated ability to localize hallucinations effectively in real-time without additional inference costs.
  • Mechanistic analysis shows improved representation of uncertainty geometry, enhancing detection reliability.
  • Achieved high performance metrics (0.966 AUROC and >80% localization accuracy) across Qwen and Llama-based models.
Read more
Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents
Ying Yuan
Large Language Models Theory Interpretability
  • Detecting an effect does not equate to learning to act on it; a reward-SNR floor governs the feasibility of learning.
  • The apparent learnability of acquisition policies is often an artifact of noise rather than a true signal.
  • Structured Hypothesis Embeddings (SHE) provide a method for generating user intent hypotheses but show limited downstream value.
  • A necessary condition for effective policy learning is that the reward SNR exceeds a defined threshold.
Read more
Procedural Fairness Failures in RLHF from Preference Averaging
M P V S Gopinadh, Karthik Kamuju, Kummari Avinash, Muppana John Joshua, Srinivasa Raju Rudraraju
Reinforcement Learning Large Language Models Optimization
  • Standard RLHF methods assume preference homogeneity, leading to procedural fairness failures.
  • Majority preference groups dominate reward learning, disadvantaging minority preferences.
  • PA-RLHF separates optimization across preference modes, preserving distinct preference signals.
  • Controlled experiments showed PA-RLHF improved alignment accuracy from 46.9% to 67.9%.
Read more
Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry
Daphne Feng, Ricardo Parada, Lily Jiang, Sophia Yi, William Chang
Theory Optimization Reinforcement Learning
  • Introduces three problem formulations for multi-agent bandits with heavy-tailed rewards and information asymmetry.
  • Develops robust decentralized algorithms for each formulation with regret guarantees that nearly match centralized rates.
  • Validates theoretical findings through experiments in a Pareto-distributed reward environment.
  • Explores trade-offs between synchronization, coordination, and exploration in decentralized settings.
Read more
Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits
Hangqi Ren, Junyi Liao
Reinforcement Learning
  • Introduction of the Counterfactual Clinical Audit (CCA) framework to evaluate medical RL agents.
  • Identification of 'Toxic Mimicry' as a critical failure mode in RL applications for sepsis management.
  • Demonstration that standard evaluation metrics fail to capture safety concerns in clinical settings.
  • Validation of CCA using the MIMIC-III database, revealing significant differences in agent behavior.
Read more
Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control
Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou
NLP Large Language Models
  • Introduction of a situated B-data framework for studying LLM behavioral personality.
  • Construction of Behavioral Mode Axes (BMAs) for effective behavioral control.
  • Demonstration of stable and model-specific behavioral profiles in LLMs.
  • Comparison of thought-derived and response-derived BMAs, highlighting the advantages of the former.
Read more
Click2Poly: A VLM for vector mapping buildings and walls
Nicolas Girard, Jawher Ben Abdallah, Arno Gobbin, Liuyun Duan, Sacha Lepretre
Computer Vision Multimodal Generative Models
  • Click2Poly enhances manual vector mapping of buildings and walls using a human-in-the-loop approach.
  • The system is built on the Florence-2 Vision Language Model, allowing for interactive editing through user clicks.
  • A comprehensive dataset was utilized, consisting of over 377,000 building polygons and 197,000 wall linestrings.
  • The implementation as a QGIS plugin facilitates real-world application in geospatial mapping.
Read more
Air Quality Station Simulation via LSTM and Attention-Based Modelling
Alexander Kostadinov, Petar O. Hristov, Dessislava Petrova-Antonova
Time Series
  • Introduction of SATADL, a deep learning model for simulating air quality station data during outages.
  • Utilizes spatial and temporal attention mechanisms to improve prediction accuracy.
  • Demonstrated superior performance in forecasting PM10 concentrations compared to baseline models.
  • Addresses a significant gap in real-time air quality data simulation and analysis.
Read more
Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization
Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes, Andrew Gordon Wilson
Theory Optimization Generative Models
  • Epiplexity is introduced as a measure of structural information in data that can enhance OOD generalization.
  • The authors propose EpiSelect and EpiGen as methods for data selection and synthetic data generation, respectively.
  • Higher epiplexity is shown to correlate with better performance in downstream tasks.
  • The paper identifies limitations in current benchmarks for evaluating data selection methods.
Read more
DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains
Shiqi Huang, Jiani He, Dingyan Shang, Yihua Xu, Jize Li, Yan Lyu, Lashimi Muraleedharan Nair
Optimization Theory Graph Learning
  • DACRI addresses the gap between disruption detection and effective intervention selection in supply chains.
  • The CriticalSCM-Bench v1 benchmark provides a controlled environment for evaluating causal interventions.
  • LambdaMART outperforms static benchmarks in certain supply chain archetypes but not universally.
  • Intervention effectiveness is influenced by timing, cost, and the nature of disruptions.
Read more
Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning
Mirko Konstantin, Stefan Zachow, Anirban Mukhopadhyay
Federated Learning
  • Introduces LIGHTYEAR, a federated learning framework that selects updates based on function-space evaluation rather than parameter-space similarity.
  • Utilizes an NTK-based agreement score to characterize predictive behavior for optimal aggregation.
  • Employs a decentralized peer-to-peer topology for direct client-to-client update exchanges, enhancing personalized aggregation.
  • Demonstrates superior performance compared to traditional centralized FL and existing P2P approaches across various datasets.
Read more
Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting
Kiran Madhusudhanan, Christian KlΓΆtergens, Lars Schmidt-Thieme, Vijaya Krishna Yalavarthi
Time Series
  • Introduces TORF, a two-stage framework for time series forecasting that separates mean prediction from uncertainty estimation.
  • Utilizes a deterministic model for accurate mean predictions in the first stage, followed by a Restricted Normalizing Flow for modeling residuals.
  • Achieves state-of-the-art performance in terms of deterministic accuracy (NMAE) and density estimation (CRPS) on various forecasting tasks.
  • Demonstrates that the two-stage approach effectively circumvents the trade-off between mean accuracy and distributional flexibility.
Read more
SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features
HyeonJun Lee, Hyeonsik Jo, Jinwoo Chung, Jangho Kim
Efficient ML Theory
  • SQuaT eliminates the lower bound on distillation loss by aligning teacher and student feature representations.
  • The framework allows for effective QAT without the need for labeled data.
  • SQuaT shows significant performance improvements in low-bit quantization settings.
  • The method is applicable across various model architectures and quantization levels.
Read more
$Ξ²$-VAEs as Effective Theories: Tolerance-Dependent Dimension
Johannes Hirn
Generative Models Theory Efficient ML
  • Increasing regularization strength in $Ξ²$-VAEs collapses low-utility latent coordinates.
  • Nonlinear interactions shift and broaden the onset of latent coordinate collapse.
  • The effective dimension is not a fixed integer but varies with reconstruction tolerance.
  • A head–tail tradeoff exists where deeper networks improve utility concentration but reduce tail fidelity.
Read more
A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression
Wouter W. L. Nuijten, Esther G. van Pelt, Albert Podusenko, İsmail Şenâz, Wouter M. Kouw
Theory Efficient ML Time Series
  • Introduces a factor graph formulation for scalable multi-output Gaussian process regression.
  • Achieves linear computational complexity in the number of data points while handling missing observations efficiently.
  • Demonstrates competitive performance against traditional kernel-matrix and inducing-point methods in empirical tests.
  • Utilizes a nearest-neighbor chain to structure inputs, facilitating exact Gaussian message passing for inference.
Read more
DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents
You Lu, Kun Zhang, Bihuan Chen, Xin Peng
Large Language Models Optimization
  • DOCSCHISEL optimizes tool documentation based on empirical analysis of LLM agent performance.
  • The effectiveness of tool documentation varies significantly across different task domains and LLM architectures.
  • The framework improves task success rates by 95.89% over original documentation and 75.15% over existing baselines.
  • The study highlights the importance of adaptive documentation for LLM agents in real-world applications.
Read more
Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
Burc Gokden
Large Language Models Theory Graph Learning
  • PLGA generalizes SDPA by using a learned bilinear operator, enhancing flexibility in attention mechanisms.
  • The architecture exhibits empirical collapse at inference, allowing for significant simplifications in output generation.
  • Operator invariance is observed, with outputs remaining stable under minor input perturbations.
  • A learned singularity condition indicates that the generator matrix becomes numerically singular at convergence.
Read more
MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale
Yuhang Yao, Zeyu Wang, Wanyi Chen, Tongyun Yang, Yuhang Han, Jie Xiao, Chengke Bao, Tianyi Zhao, Lynn Ai, Eric Yang, Tianyu Shi
Large Language Models NLP Efficient ML
  • MERA improves small model capabilities through iterative adaptation rather than just routing.
  • The framework utilizes a SkillBook to capture and distill successful execution patterns.
  • Empirical results show significant performance gains on benchmark tasks.
  • Verifier-backed fallback mechanisms ensure quality during deployment.
Read more
Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning
Vaneet Aggarwal
Optimization Theory Efficient ML
  • Establishes the resilience of the SGS-Poisson algorithm under controlled oracle errors.
  • Achieves classical approximation factors for both monotone and non-monotone objectives in full-bandit settings.
  • Introduces novel technical results, including adaptive potential-preservation and robust swap lemmas.
  • Demonstrates the applicability of offline optimization techniques to online learning frameworks.
Read more
High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving
Muhammad Ayub Sabir, Junbiao Pang, Fatima Ashraf
Robotics Time Series Interpretability
  • Introduction of a causal high-order liquid evidence framework for GNSS spoofing detection.
  • Utilization of physics-guided residual evidence to model GNSS-motion inconsistencies.
  • Separate adaptive liquid encoders for processing different orders of evidence variations.
  • Achieved highest F1-scores among evaluated models on real-world datasets.
Read more
Retrieval-Corrected Conformal Prediction for Time Series
Sangjin Jin, Kangmin Kim, Junhyeong Lee, Yongjae Lee
Time Series
  • RCCP improves upon traditional conformal prediction methods by using retrieval of similar past residuals for local calibration.
  • The method constructs asymmetric prediction intervals that reflect local error behavior without relying on global residual quantiles.
  • RCCP achieves target coverage levels and lower Winkler scores across various benchmarks, indicating improved performance.
  • The approach maintains low calibration and inference overhead, enhancing its scalability for practical applications.
Read more
Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches
Muntasir Hasan Kanchan, Md. Alamgir Hossain, Md. Samiul Islam, Muhammad Masud Tarek
NLP
  • Introduces a dual-model framework combining machine learning and deep learning for sentiment analysis.
  • Evaluates five machine learning algorithms and five deep learning models on an imbalanced dataset.
  • Implements a preprocessing pipeline using NLP techniques to enhance input quality.
  • Demonstrates model performance on unseen real-world data, reflecting practical deployment scenarios.
Read more
FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation
Yu Zhang, Zhihan Wang, Guanlin Chen, Min Jiang, Shuai Li
Optimization Theory
  • FunnelCausalNet integrates conversion and revenue uplift modeling, addressing the limitations of decoupled approaches.
  • The model incorporates a variance analysis that guides its implementation under high zero-inflation regimes.
  • A budgeted multi-tier allocation strategy is proposed, allowing for efficient resource distribution across different coupon strengths.
  • Empirical evaluations show significant improvements in GMV effect error reduction and ROI metrics compared to existing baselines.
Read more
Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting
Junyi Ye, Ivy Gateri Wanjiku
Time Series
  • First systematic study of activation calibration for PTQ in financial forecasting.
  • 4-bit activation quantization leads to significant predictive losses, recoverable through improved calibration.
  • Activation range preferences evolve over time, necessitating dynamic calibration strategies.
  • Practical deployment guidelines are provided for selecting among various quantization methods.
Read more
Long-Time Trajectory Approximation via SA-NODEs: Model Predictive and Floquet Strategies
Ziqian Li, Nikolaos M. Matzakos
Theory Optimization Time Series
  • Introduces two strategies to mitigate error growth in long-time trajectory approximation using SA-NODEs.
  • The model predictive strategy utilizes adaptive partitioning and state resets to maintain error tolerance.
  • The Floquet strategy leverages stable limit cycles to ensure linear error growth without requiring data at deployment.
  • Numerical experiments confirm the effectiveness of the proposed strategies and their theoretical guarantees.
Read more
A Recommendation System Approach for Interference-Robust Sensor Subset Selection
Kaan Buyukkalayci, Kyle Pak, Merve Karakas, Christina Fragouli
Audio & Speech Efficient ML Robotics
  • Formulates sensor activation as a recommendation problem using network observations as context.
  • Introduces a Two-Tower MLP architecture for efficient scoring of sensor subsets.
  • Demonstrates significant improvements in robustness to acoustic interference using frequency-band features.
  • Achieves high accuracy (up to 98.4%) while maintaining low computational requirements (sub-millisecond to 1 ms).
Read more
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
Reinforcement Learning Large Language Models NLP
  • Rubric-as-reward RL can lead to reward hacking due to fixed criteria exploitation.
  • Rubric Dropout is a simple yet effective method to mitigate this issue by randomly dropping rubric criteria during training.
  • Experiments show significant improvements in OOD performance and reductions in reward hacking metrics.
  • The method is computationally efficient, requiring only a single hyperparameter adjustment.
Read more
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
Xinmu Ge, Zizhuo Zhang, Yu Huang, Jianing Zhu, Lin Yuan, Wanli Gu, Weichang Wu, Weiran Huang, Xiaolu Zhang, Bo Han, Jun Zhou, Jiangchao Yao
NLP Large Language Models Theory
  • OPD improves sampling efficiency but does not consistently expand the reasoning capability boundary of student models.
  • OPD-trained models show better performance at small K values but are surpassed by pre-OPD models at larger K.
  • More previously solvable problems become unsolvable after OPD training than vice versa.
  • The study introduces the concept of 'illusory distillation,' indicating that apparent gains stem from better access to existing capabilities.
Read more