AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

62 Papers today
8h Update frequency
7 Days of history
Muse: Representation Geometry of Muon Beyond Normalized Momentum
Da Chang, Qiankun Shi, Lvgang Zhang, Di He, Yaoshuai Ma, Ganzhao Yuan, Yongxiang Liu
Optimization Large Language Models Theory
  • Muse optimizers leverage different matrix representations to enhance Muon-style optimization.
  • Each representation induces distinct polar geometries affecting convergence and performance.
  • Balanced non-native representations can match the performance of native representations.
  • Reducing the shorter dimension in representations weakens optimizer performance.
Read more
Revisiting data-driven dynamic security assessment with a tabular foundation model
Olayiwola Arowolo, Maosheng Yang, Jochen Cremer
Theory Optimization Efficient ML
  • Introduction of a tabular foundation model (TFM) for dynamic security assessment (DSA) in power systems.
  • Elimination of the need for separate models for each contingency, enhancing efficiency.
  • Demonstration of improved generalization to unseen contingencies using electrical distance coordinates (EDC).
  • Achievement of high accuracy with significantly fewer labeled samples compared to traditional methods.
Read more
Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment
Xinran Liu, Yuwen Li, Hongxiang Gao, Heyang Xu, Jianqing Li, Zongmin Wang, Chengyu Liu
Time Series Multimodal
  • PEACE framework effectively transfers ECG interpretation from adult to pediatric populations using structured clinical knowledge.
  • Label-conditioned representation and alignment enhance the model's ability to utilize diagnostic descriptors during training.
  • Curriculum adaptive fusion optimizes alignment strength based on training progress, improving performance under limited pediatric supervision.
  • The model achieves substantial gains in AUC scores, outperforming traditional domain adaptation methods.
Read more
Information-Directed Sampling for Causal Bandits
Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi
Theory Optimization
  • Introduces a Bayesian framework for contextual causal bandits with non-manipulable variables.
  • Develops causal variants of Thompson Sampling and Information-Directed Sampling.
  • Establishes entropy-dependent regret bounds for both methods.
  • Demonstrates the effectiveness of the proposed methods through experiments on synthetic tasks.
Read more
Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints
Patrick Inoue, Florian Röhrbein, Andreas Knoblauch
Multimodal Efficient ML Theory
  • Hebbian learning can optimize synaptic resource allocation under structural constraints.
  • The study introduces a mutual-information-based metric for evaluating representational cost.
  • Hebbian learning outperforms sparse BP and DDTP in terms of task-information cost.
  • The results highlight a cost-performance trade-off rather than uniform accuracy improvements.
Read more
Learning Faster without Deeper Networks: A*-Inspired Batch Selection for Efficient CNN Training
Anxhelo Shehu, Enes Stastoli, Arben Cela
Efficient ML
  • A*-Inspired Batch Selection (A*-BS) improves CNN training efficiency by optimizing mini-batch selection.
  • The method combines a loss-based difficulty measure with a reuse penalty to rank batches.
  • A*-BS outperforms random batch shuffling in all evaluated tasks and achieves higher accuracy than deeper models.
  • The approach is lightweight, model-agnostic, and can be integrated into existing training pipelines.
Read more
When Does Muon Help Agentic Reinforcement Learning?
Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
Reinforcement Learning Optimization
  • Muon optimizer shows an 88% improvement in validation success over AdamW in sparse-reward RL tasks.
  • The effectiveness of Muon is dependent on the choice of advantage estimator and learning rate.
  • Applying Muon only to hidden weight matrices yields significant performance gains in GiGPO.
  • Learning-rate controls demonstrate that increasing AdamW's rate does not replicate Muon's success.
Read more
Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation Forecasting
Akshay Sunil, Muhammed Rashid, Raja Sekhar Sivaraju, Sushma Nair, Subimal Ghosh
Time Series
  • Developed a radar-only nowcasting framework using a Multi-Variable U-Net for high-resolution precipitation forecasting.
  • Incorporated multi-elevation reflectivity and Doppler velocity data to enhance prediction accuracy.
  • Achieved significant improvements in forecasting skill compared to traditional persistence methods.
  • Utilized a high-reflectivity attention module to focus on convective cores and ensure meteorological relevance.
Read more
Field-Aware RankMixer with Dual-Stream Bilinear Fusion for the Tencent UNI-REC Challenge
Yufeng Zhang, Zhengqi Xu, Jiajun Cui
Optimization Multimodal
  • Development of the FA-RankMixer model for unified feature modeling in pCVR prediction.
  • Utilization of target-aware DIN modules to extract user interests from multiple behavior domains.
  • Implementation of dual-stream bilinear fusion to enhance representation learning.
  • Achieved ninth place in the Tencent UNI-REC Challenge, indicating practical applicability.
Read more
Improving Improved Kernel PLS
Ole-Christian Galbo Engstrøm
Efficient ML Theory Optimization
  • Introduces optimizations to the IKPLS algorithms for faster computation.
  • Replaces term-by-term accumulation in R computation with a direct evaluation strategy.
  • Identifies new mathematical equivalences for Y loading computations, reducing their cost.
  • Demonstrates significant speed improvements in benchmarks on both CPU and GPU.
Read more
Kolmogorov--Arnold Networks for Small Language Models
Felippe Alves, Renato Vicente
NLP Large Language Models Interpretability
  • KANs provide a practical interface for auditing learned transformations in small language models.
  • Interpretability is enhanced as KAN edge functions can be reconstructed and analyzed exhaustively.
  • No consistent benchmark advantage was found for KANs over strong MLP baselines on standardized tests.
  • Pruning low-activity edge functions in KANs is effective but not unique compared to MLPs.
Read more
A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning
Yang Liu, Yuhao Liu, Yunran Wei
Reinforcement Learning Optimization Theory
  • Integration of IRL and RL for risk preference elicitation and optimization.
  • Adaptive Bayesian IRL method allows for the inference of risk objectives from noisy decisions.
  • Model-free RL algorithm optimizes policies under a broad class of distortion riskmetrics.
  • Quantile neural networks enable the estimation of conditional cost quantile functions.
Read more
In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention
Katsuyuki Hagiwara
Theory Optimization
  • Introduces a transformer model that learns the closed-form solution for simple linear regression.
  • Utilizes layer normalization to approximate division required for the least squares estimate.
  • Demonstrates that the transformer can effectively learn the analytical solution rather than relying on gradient descent.
  • Presents numerical experiments validating the model's performance under â„“1 regularization.
Read more
Low-Latency Relay Selection in NR-V2X Vehicular Communications via Graph Isomorphism Networks with Edge Features
Giambattista Amati, Federica Mangiatordi, Emiliano Pallotti, Simone Angelini, Pierpaolo Salvo, Paola Vocca
Graph Learning Optimization
  • Introduces an edge-aware Learning-to-Optimize framework for relay selection in NR-V2X networks.
  • Utilizes Graph Isomorphism Networks with Edge Features to model V2X scenarios as directed graphs.
  • Achieves high accuracy and F1-scores in relay selection while maintaining low inference latency.
  • Presents a hybrid GINE-Pruned MILP strategy that preserves optimality and reduces computation time.
Read more
ChronoQG: Towards a Temporally Expressive and Hop-Bounded Benchmark for Temporal Knowledge Graph Question Generation
Xuemeng Liu, Zhengpin Li, Wanpeng Tang, Haotong Xie, Wentao Zhang
NLP Graph Learning Large Language Models
  • ChronoQG is the first benchmark specifically designed for Temporal Knowledge Graph Question Generation (TKGQG).
  • The framework incorporates a comprehensive taxonomy of temporal constraints and topology-temporal subgraph sampling.
  • The study reveals that existing KGQG methods have difficulty preserving temporal constraints, especially in multi-constraint scenarios.
  • ChronoQG produces a total of 16,011 verified questions from heterogeneous temporal knowledge graphs.
Read more
Auditing Fairness-Privacy Trade-offs: Subpopulation-Level Effects of Fairness-Enhancing Algorithms
Umid Suleymanov, Ilhama Novruzova, Khalid Mammadov, Natavan Hasanova, Murat Kantarcioglu
Theory
  • First systematic audit of fairness-enhancing algorithms' effects on subpopulation-level privacy risks.
  • Extension of the Likelihood Ratio Attack (LiRA) for subgroup-specific privacy auditing.
  • Differential Privacy's interaction with fairness interventions reveals uneven privacy-utility trade-offs.
  • Fairness and privacy are not inherently conflicting; their relationship is influenced by model architecture and subgroup representation.
Read more
Data-Native Global Optimization for Big Data K-means Clustering
Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan
Optimization Efficient ML
  • Introduction of Big-means++, a scalable algorithm for K-means clustering on big data.
  • Utilizes random samples to create surrogate landscapes for global optimization.
  • Incorporates a flowing-incumbent strategy to enhance centroid mobility.
  • Employs a shaking mechanism to vary sample sizes and improve solution quality.
Read more
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian Ramamoorthy
Reinforcement Learning Robotics Efficient ML
  • RENEW addresses model exploitation in offline RL by using human preferences instead of expert demonstrations.
  • The proposed DLHF framework formalizes the learning of dynamics from human feedback.
  • RENEW improves sample efficiency and reduces catastrophic forgetting compared to naive DLHF.
  • The method effectively targets transitions where the model is most uncertain, enhancing model reliability.
Read more
DebrisTracer: Reliable Tracking in Hypervelocity Impact Fast Imaging
Théophane Loloum, Fabien Vivodtzev, David Hébert, Baptiste Reynier, Michel Arrigoni, Julien Tierny
Computer Vision
  • DebrisTracer improves tracking accuracy of debris in hypervelocity impact imaging.
  • The framework incorporates domain knowledge and physical assumptions for better reliability.
  • Extensive experiments show significant accuracy improvements over existing tools.
  • The approach enables interpretable visual analyses of debris dynamics.
Read more
Who Became Financially Vulnerable After COVID-19? A Population-Level Machine Learning Analysis Using MEPS Data
Alexey Kresin, Zien Cheng, Ammar Ahad, Ebiyomare Kelvin, Manish Sivaratri, Prabhjeet Singh, Omar Aljawfi, Olabisi Ojo, Nawar Shara
Interpretability
  • The study analyzes financial vulnerability in healthcare pre- and post-COVID-19 using MEPS data.
  • High financial burden is defined as out-of-pocket healthcare spending greater than 10% of family income.
  • Income level, insurance status, and prescription drug spending are significant factors influencing financial vulnerability.
  • Subgroup analyses reveal persistent disparities in financial burden across different demographic groups.
Read more
Explainable Geospatial AI for Satellite Ground Station Siting Using LiDAR-Derived Terrain Intelligence
Shohini Sarkar, Smithi Mahendran, Rishi Chudasama, Varun Mannam, Arav Luthra, Yuvraj Rekhi, Vivek Nadig, Arsh Goenka
Interpretability
  • Introduces a machine learning framework for predicting Representative Clutter Height (RCH) using open geospatial data.
  • Achieves significant accuracy improvements over traditional fixed clutter height methods, reducing MAE by over 60%.
  • Utilizes LightGBM for its efficiency and compatibility with feature attribution analysis.
  • Identifies influential predictors for RCH, enhancing interpretability and transparency of the model.
Read more
CoG-Guided Weight Correction for Fault-Tolerant Deep Neural Networks
Bahram Parchekani, Samira Nazari, Ali Azarpeyvand, Mohammad Hasan Ahmadilivani, Tara Ghasempouri, Jaan Raik
Time Series Theory Efficient ML
  • Introduces a CoG-guided weight correction method for DNNs.
  • Eliminates the need for retraining or architectural changes.
  • Demonstrates significant fault tolerance improvements in LSTM and CNN architectures.
  • Validates the approach using real-world datasets in safety-critical applications.
Read more
Recursive Harness Self-Improvement
Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, Yujin Tang
Optimization Robotics Efficient ML
  • RHI optimizes user-constructed harnesses to improve execution trace quality.
  • The approach is computationally lightweight, requiring only a few iterations for significant performance gains.
  • RHI outperforms high-reasoning-effort agents while reducing inference costs by up to 60%.
  • Improvements stem from enhanced task-specific context management and inter-agent information flow.
Read more
PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
Yuchen Yang, Yifan Zhao, Anisha Dasgupta, Sasa Misailovic
Large Language Models Efficient ML NLP
  • PagedWeight dynamically quantizes MoE model weights at runtime, optimizing GPU memory usage.
  • Achieves significant memory savings (up to 72.0%) and throughput improvements (1.94×) compared to existing methods.
  • Implements a quality-aware runtime planner that minimizes the impact of quantization on task accuracy.
  • Utilizes asynchronous page movement to reduce latency during model serving.
Read more
Diffusion models recover accurate mixture weights despite score function insensitivity
Andrew Dennehy, Ramchandran Muthukumar, Rebecca Willett, Nisha Chandramoorthy
Generative Models Theory
  • Introduces the diffusion score sensitivity index (DSSI) to quantify the sensitivity of the DSM loss to parameter changes.
  • Proves that mixture weight estimation errors are on the same order as the DSM loss for Gaussian mixtures.
  • Demonstrates that the choice of noise schedule can significantly affect the sensitivity of diffusion models.
  • Empirical results show that DSSI can predict the accuracy of mixture weight recovery in trained models.
Read more
An Exploratory Study of Single Channel Surface Electromyography for Hand Gesture Classification
Daanish Hindustani
Efficient ML
  • Single-channel sEMG can effectively classify hand gestures, challenging the reliance on multichannel systems.
  • Feature engineering techniques, including Pearson correlation filtering, significantly enhance classification performance.
  • A compact neural network can achieve competitive accuracy (up to 90%) with limited data input.
  • The study supports the feasibility of deploying sEMG systems in low-power environments like Raspberry Pi.
Read more
On-Policy Delta Distillation
Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
Reinforcement Learning Large Language Models NLP
  • Introduction of On-Policy Delta Distillation (OPD2) as an improvement over traditional on-policy distillation methods.
  • The delta signal, defined as the difference between the teacher model and its base model, serves as a more effective distillation reward.
  • Extensive empirical validation shows OPD2 outperforms conventional methods in reasoning tasks across multiple domains.
  • The paper proposes two reward design strategies to enhance the effectiveness of the delta signal in the distillation process.
Read more
When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
S. Aaron McClendon
Reinforcement Learning Theory Optimization
  • Model merging can match the performance of joint multi-task RL when data can be pooled.
  • Merging methods (TIES, RAM+) yield statistically indistinguishable results from joint training.
  • Task vectors of specialists are near-orthogonal, indicating minimal interference during merging.
  • A calibrated geometric analysis helps explain the merging outcomes and their effectiveness.
Read more
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Weiwen Xu, Jia Liu, Hou Pong Chan, Long Li, Deng Cai, Min Chen, Hao Zhang
Reinforcement Learning Large Language Models Optimization
  • CPO introduces contrastive disagreement as a more reliable correctness signal compared to entropy.
  • The framework effectively addresses the zero-advantage problem in RLVR.
  • CPO demonstrates substantial performance improvements over existing entropy-based methods.
  • The study provides a theoretical foundation for on-policy distillation (OPD) as a special case of CPO.
Read more
BadWAM: When World-Action Models Dream Right but Act Wrong
Qi Li, Xingyi Yang, Xinchao Wang
Robotics
  • Introduction of World-Action Drift Attacks as a new vulnerability in WAMs.
  • Development of BadWAM framework to evaluate and model these attacks.
  • Demonstration of significant task performance degradation due to adversarial perturbations.
  • Identification of a critical gap between imagined futures and executed actions in WAMs.
Read more
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Jan Betley, Johannes Treutlein, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans
NLP Large Language Models
  • Covert value leakage occurs when LLMs' answers are influenced by their own values without disclosure.
  • Models exhibit biases based on moral outcomes, company affiliation, and preferences for leisure activities.
  • Significant differences in value leakage were observed among various frontier models.
  • Current alignment training does not adequately address the issue of value leakage.
Read more
On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures
Mohamed Amine Kina
Generative Models Theory Optimization
  • CAKE's boundary-seeking principle is ill-suited for autoencoder architectures due to their shared latent manifold.
  • Independent sampling of pixel-level targets leads to severe gradient conflicts in autoencoders.
  • A single noise forward pass is proposed as a more effective baseline for data-free generative distillation.
  • The study reformulates autoencoder tasks into a classification framework to facilitate comparison with CAKE.
Read more
Scalable Training of Continuous-Time Spiking Neural Networks with Differentiable Spike-Time Discretization
Yusuke Sakemi, Tomoya Takeuchi, Takeo Hosomi, Kazuyuki Aihara
Efficient ML Theory Time Series
  • Introduces a memory-efficient framework for training continuous-time SNNs using differentiable spike-time discretization (DSTD).
  • Reduces memory requirements for candidate spike times significantly, enabling training of deeper networks.
  • Incorporates temporal regularization to mitigate dead-neuron issues and enhance processing efficiency.
  • Achieves substantial reductions in peak memory consumption and training time compared to traditional methods.
Read more
Structure of the Circular-Dyadic Convolution Error
Ben Fauber, Alireza Moradzadeh
Theory Efficient ML
  • Exact error cancellation occurs at specific input and output positions.
  • The error operator is nearly full rank with a logarithmic null space dimension.
  • A single alignment scalar governs the expected substitution error.
  • Substitution error typically doubles output energy, except for filters in a zero-error subspace.
Read more
Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems
Matteo Tomasetto, Nicolò Botteghi, Gabriele Bruni, Andrea Manzoni
Reinforcement Learning Optimization Robotics
  • PEARL integrates traditional optimal control with reinforcement learning to enhance sample efficiency.
  • The actor-adjoint algorithm reduces the number of environment interactions needed for training.
  • PEARL demonstrates superior performance in high-dimensional dynamical systems compared to existing RL methods.
  • The approach generalizes well across multiple scenarios, crucial for parametric systems.
Read more
A Transportable Threshold-Based Framework for Interpretable Classification of Medical Data
Antony Garcia, Adrian Noriega, Gabrielle Britton, Xinming Huang
Interpretability
  • Introduces a framework for interpretable classification using Bernoulli Naïve Bayes.
  • Utilizes χ2-guided statistical binarization for continuous medical data.
  • Achieves high AUC scores on benchmark medical datasets.
  • Demonstrates improved probability calibration through rigorous analysis.
Read more
AI Trading: Evaluating Large Language Models for Technical Market Analysis
Geofrey Ntale
NLP Large Language Models Time Series
  • Systematic evaluation of five LLMs for technical market analysis.
  • GPT-4 Turbo achieves the highest annualized return and Sharpe ratio among general-purpose models.
  • FinGPT demonstrates competitive performance due to domain-specific fine-tuning.
  • Common failure modes include numerical hallucination and inconsistent performance in sideways markets.
Read more
GAttNHP: Group Attention Neural Hawkes Process for Extrapolation Reasoning in Temporal Knowledge Graphs
Xiangni Tian, Kaixian Yu, Runpeng Dai, Niansheng Tang, Hongtu Zhu
Graph Learning Time Series
  • GAttNHP effectively captures long-range temporal dependencies in TKGs.
  • The framework allows for mutual excitation between event chains through a soft-grouping mechanism.
  • Non-Crossing Quantile regression provides robust time predictions, addressing issues with heavy-tailed distributions.
  • GAttNHP outperforms existing models on benchmark datasets, particularly in challenging long-tail scenarios.
Read more
Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling
Moxian Qian
Theory Generative Models Optimization
  • NHMC combines global proposal moves with statistical correction for unnormalized Boltzmann sampling.
  • The method utilizes non-equilibrium work identities to correct sampling estimates.
  • Training involves minimizing mean generalized work to enhance path-space sampling accuracy.
  • NHMC shows effective performance on specific target distributions while highlighting limitations in path overlap scenarios.
Read more
Angular Gaussian Supervised Contrastive Learning for Long-Tailed Electrocardiogram Arrhythmia Diagnosis
Jin Dai, Qiuzhen Zhang, Chenyun Dai, Danmei Lan, Can Han
Time Series
  • AG-SCL improves ECG arrhythmia diagnosis in long-tailed distributions.
  • The method combines direction-aware contrastive learning with adaptive logit adjustments.
  • Significant performance gains were observed in rare arrhythmia categories.
  • The approach preserves critical ECG morphology during data augmentation.
Read more
Learning Who to Treat When Treatment is Missing
Johnna Sundberg, Rayid Ghani, Eli Ben-Michael, Edward Kennedy
Theory Efficient ML Optimization
  • Introduces efficient estimators for CATE and policy value under missing treatment data scenarios.
  • Proves that MAR estimators are more efficient than MCCAR estimators, providing a formal argument against complete-case analysis.
  • Empirical validation shows that correctly specifying the missingness mechanism is crucial for accurate estimation.
  • Offers theoretically grounded tools for robust policy learning in the presence of missing treatment data.
Read more
Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework
Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
Graph Learning Federated Learning Multimodal
  • Introduces FedGAMMA, a novel framework for federated multimodal graph learning.
  • Addresses critical challenges in privacy, semantic-structural alignment, and client heterogeneity.
  • Achieves state-of-the-art performance across multiple multimodal graph datasets.
  • Utilizes a two-stage approach combining pre-training and prompt-based fine-tuning.
Read more
NeuroGRIP: Retrieval-Augmented Graph Refinement for Knowledge-Grounded EEG Seizure Diagnosis
Lincan Li, Zheng Chen, Yushun Dong
Graph Learning Time Series Interpretability
  • NeuroGRIP integrates external medical knowledge into EEG seizure diagnosis using a retrieval-augmented graph refinement approach.
  • The framework constructs a domain-specific knowledge base from clinical guidelines to enhance the interpretability of EEG graphs.
  • Semantic alignment queries are used to match EEG node embeddings with clinical relations, improving the robustness of predictions.
  • Extensive experiments show significant improvements in seizure detection accuracy and interpretability compared to traditional STGNN methods.
Read more
Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap
Olivier Jeunen
Theory Efficient ML
  • Introduces a novel experimental protocol leveraging policy overlap to reduce variance in A/B-testing.
  • Proves that the new estimators dominate the standard Difference-in-Means estimator under certain conditions.
  • Identifies optimal traffic allocation strategies based on policy divergence.
  • Develops the Δ-MRDR estimator to minimize variance in ATE estimation.
Read more
HyperShadow: A Benchmark for Detecting 3D Projections of Higher-Dimensional Spatial Objects
Akshay Sasi
Computer Vision Theory
  • HyperShadow is the first benchmark for detecting 3D projections of higher-dimensional objects.
  • Standard intrinsic-dimension estimators fail to accurately identify shadows, achieving only 71-73% accuracy.
  • A compact learned point network achieves 96.2% accuracy in shadow detection.
  • The introduced rigidity witness statistic effectively distinguishes between rigid and non-rigid motions.
Read more
Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data
Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
Generative Models Time Series
  • Introduces a taxonomy-guided evaluation protocol for assessing temporal fidelity in synthetic sequential tabular data.
  • Develops Seq2Synth, a benchmark that measures timestamp validity, cross-sectional fidelity, longitudinal fidelity, and structural consistency.
  • Demonstrates that conventional evaluation methods can misrepresent the performance of generative models by ignoring temporal dynamics.
  • Finds substantial differences in rankings of generative models when evaluated temporally versus statically.
Read more
CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach
Andrei Neagu, Eeham Khan, Leila Kosseim
Reinforcement Learning Time Series NLP
  • Integration of multi-modal features (technical indicators, sentiment scores, cyclical encodings) enhances trading decision-making.
  • Introduction of an alpha reward mechanism aligns training objectives with the goal of outperforming a buy-and-hold strategy.
  • DDPG demonstrates the best performance across both BTC and TSLA, significantly exceeding the buy-and-hold baseline.
  • The study reveals challenges in adapting trading policies to different market conditions, particularly between bull and bear markets.
Read more
PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference
Hyunwoo Oh, Suyeon Jang, Hanning Chen, KyungIn Nam, Sanggeon Yun, Ryozo Masukawa, Mohsen Imani
Large Language Models Efficient ML
  • PolyQ enables flexible, activation-aware channel-wise quantization for CPUs, enhancing deployment efficiency.
  • The framework reduces activation reorder traffic by up to 70.8% through compile-time layout regularization.
  • PolyQ achieves significant perplexity improvements (2.4–32.1%) at a 3-bit target compared to prior methods.
  • End-to-end measurements show that PolyQ maintains low energy overhead (<2%) and latency (within 5.8%) relative to optimized backends.
Read more
Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes
Jaeyeong Lee, Taeseong Yoon, Wonmo Koo, Heeyoung Kim
Graph Learning Time Series
  • Introduces a knowledge-assisted multi-graph framework for anomaly detection in industrial processes.
  • Constructs three complementary graphs to model sensor dependencies, enhancing anomaly detection performance.
  • Utilizes a multi-graph attention network for improved representation of complex dependencies.
  • Demonstrates significant performance improvements in anomaly detection through experiments on real-world datasets.
Read more
LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models
Jingteng Li, Alexander Capstick, Louise Rigny, Iona Biggart, Neil J Sebire, Payam Barnaghi
Multimodal Large Language Models Time Series
  • LLM4EHR introduces a multimodal framework for aligning EHR event sequences with clinical time series.
  • The model achieves improved performance on key clinical prediction tasks compared to existing methods.
  • LLM4EHR demonstrates strong transferability of learned embeddings to new cohorts.
  • The semantic regularised contrastive objective is essential for enhancing prediction outcomes.
Read more
Graph Coloring Approach to Solving Sudoku with Oscillatory Neural Networks
Filip Sabo, Aida Todri-Sanial
Optimization
  • Introduction of a novel ONN-based solver for Sudoku using Graph Coloring.
  • Modification of existing Graph Coloring solvers to improve computational efficiency.
  • Achieved nearly flawless accuracy on 4x4 Sudoku puzzles and high accuracy on 9x9 puzzles.
  • Demonstrated significant performance improvement over existing HNN and ONN solvers.
Read more
Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation
Zhaoyang Jiang, Zhizhong Fu, Zicheng Li, Yunsoo Kim, Jiacong Mi, Xuanqi Peng, Fei Teng, Honghan Wu
NLP Large Language Models Graph Learning
  • The distinction between presentation and mechanism is crucial in evaluating memory architectures.
  • Fine-grained memory mechanisms do not significantly outperform simpler methods when presentation is controlled.
  • Coarse invalidation mechanisms are effective for current-state queries.
  • Memory evaluations should maintain a fixed render to avoid confounding results.
Read more
Deep Learning Approaches for Sleep Apnea Classification from Polysomnographic EEG Signals
Shashank Manjunath, Mukesh Cheemakurthi, Aarti Sathyanarayana
Time Series Graph Learning
  • First comprehensive comparison of multiple deep learning architectures for OSA detection from multichannel EEG.
  • Introduction of TDA-based features that effectively capture apnea signatures.
  • Significant performance variations observed across patient demographics and sleep stages.
  • Demonstrates the feasibility of automated OSA screening using EEG signals.
Read more
LIGO-PINN: Learned Initialization via Gated Optimization to Alleviate Convergence Failures in Physics Informed Neural Networks
Nilay Anurag, Shital Adhikari, Taniya Kapoor, Nikhil Muralidhar
Optimization Theory
  • LIGO-PINN introduces a novel learned initialization method to improve PINN training dynamics.
  • The gated layerwise optimization (GLO) procedure effectively mitigates convergence failures in challenging PDEs.
  • The proposed method outperforms state-of-the-art techniques, achieving significant performance improvements.
  • LIGO-PINN demonstrates generalization capabilities to 3D unstructured domains.
Read more
Decoding Market Emotion from Blockchain Activity: A Data-Driven Sentiment Classifier
Arthur G. Bubolz, Abreu Quevedo, Giancarlo Lucca, Rafael A. Berri, Eduardo Borges, Bruno L. Dalmazo
NLP Time Series Interpretability
  • Integration of blockchain data with social media sentiment to explain market behavior.
  • Focus on understanding market sentiment rather than predicting prices.
  • Gradient Boosting (XGBoost) achieved an average F1-score of 0.84 for sentiment classification.
  • Use of SHAP values for model interpretability, enhancing transparency of results.
Read more
Integration Matters: Rollout-Based Training for Constrained Diffusion Models
Xiaoxuan Liang, Saeid Naderiparizi, Berend Zwartsenberg, Frank Wood
Generative Models Robotics Optimization
  • Introduces a fine-tuning framework for constrained diffusion models that aligns training with sampling.
  • Utilizes online rollouts to measure constraint violations during the denoising trajectory.
  • Maintains distributional fidelity by incorporating a regularization term from the existing denoising loss.
  • Demonstrates effectiveness on constrained generation tasks with improved constraint satisfaction.
Read more
Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation
Himel Dev, Madhusudan Basak, Tanmoy Sen, Paromita Shome, Bashima Islam
Optimization
  • Introduces a symbolic reasoning engine achieving 99.7% candidate recall for packing checklists.
  • Develops a two-stage preference learning model that mitigates survivorship bias in user feedback.
  • Utilizes a CP-SAT optimizer to ensure 100% feasibility under complex constraints, outperforming greedy and random methods.
  • Demonstrates significant improvements in checklist completion rates and efficiency in a real-world application.
Read more
From hyperplanes to hyperellipsoids: characterizing the inherent interpretability of linear and single-qubit mixed-state binary classification models
Kaitlin Gili
Theory Interpretability
  • Single-qubit mixed-state models are analogous to linear models but learn hyperellipsoids instead of hyperplanes.
  • Linear models provide absolute feature importance, while single-qubit models offer relative feature importance.
  • The paper aims to bridge the gap between classical machine learning and quantum machine learning for educational purposes.
  • Understanding the geometric inductive biases of models can enhance model design and empirical verification.
Read more
Probabilistic Physics-Informed Neural Networks for Estimating Heterogeneous Elastic Properties from Low-Resolution and Noisy Displacement Data
Tatthapong Srikitrungruang, Jaesung Lee
Theory
  • PIE-PINN effectively estimates heterogeneous elastic properties from low-resolution and noisy data.
  • The framework combines B-spline representation for global displacement and a hierarchical model for local variations.
  • An alternating maximum-likelihood training strategy enhances robustness by dynamically adjusting loss weights.
  • Case studies confirm the robustness of PIE-PINN across different noise levels and resolutions.
Read more
Stochastic Reset Pathfinding: Path-Level Regret for Cascading Bandits over Graph Paths
Guni Sharon, Wei Zhang
Reinforcement Learning Graph Learning Theory
  • Introduction of Stochastic Reset Pathfinding (SRP) as a new episodic learning problem.
  • Demonstration of the open-loop optimal policy structure, linking SRP to the combinatorial cascading bandit framework.
  • Development of the Log-Dijkstra meta-algorithm with PathUCB and PathTS instantiations.
  • Establishment of a path-level regret bound that provides deeper insights into structured graphs.
Read more
(MPO)$^2$: Multivariate Polynomial Optimization based on Matrix Product Operators
Niccolò Ciolli, Anders Vestergaard Nørskov, Michael Kastoryano, Petr Taborsky, Morten Mørup
Optimization Theory Efficient ML
  • Introduces (MPO)² framework for multivariate polynomial optimization.
  • Combines MPO feature embeddings with compact polynomial weight tensors.
  • Offers feature order independence and incorporates structured operators.
  • Demonstrates improved performance over existing polynomial models.
Read more
TIDE: Trustworthy and Interpretable Battery Degradation Estimation with Contextual Learning and Symbolic Distillation
Wen Yang Tan, Jiawei Li, Fang Liu, Wei Zhang, Sumei Sun, Peng Cheng Wang, Elisa Y. M. Ang
Interpretability Time Series Efficient ML
  • TIDE combines accuracy, trustworthiness, and interpretability in battery health estimation.
  • The model features a three-component backbone: knowledge-guided prior, monotone residual, and contextual learning.
  • TIDE achieves an average accuracy improvement of 19.7% over baseline models.
  • The symbolic distillation process provides a compact and interpretable model representation.
Read more