AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

24 Papers today
8h Update frequency
7 Days of history
From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good
Nitesh V. Chawla, Paulo Benanti
Theory
  • AI serves as both a diagnostic tool and an intervention in existing institutional failures.
  • A 'rupture test' is proposed to evaluate the impact of AI on human relationships and institutional weaknesses.
  • The paper critiques the transition from ethical principles to protocols in AI governance.
  • The RISE AI framework is introduced to make explicit and evidence-based claims about AI's societal impact.
Read more
Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
Anqi Peter Li, Kaden Kim
Reinforcement Learning Robotics Optimization
  • Introduction of the fork ledger for evaluating individual world-model updates.
  • Demonstrated that fixed update mechanisms can negatively impact performance across multiple tasks.
  • Counterfactual comparisons reveal that the value of updates varies with task and environmental drift.
  • The methodology allows for observable counterfactual utility in simulation, enhancing decision-making processes.
Read more
General Quantification of Covariate and Concept Shifts
Hongbo Chen, Li Charlie Xia
Theory
  • Introduces ฮณโˆ—-concept shifts to address ill-defined concept shifts due to support mismatch.
  • Derives a general error bound that unifies covariate and ฮณโˆ—-concept shifts applicable to various learning tasks.
  • Proposes two estimators for accurately estimating shifts from finite samples.
  • Develops the DataShifts algorithm for practical quantification of distribution shifts.
Read more
TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription
Akshaj Gupta, Hwi Joo Park, Andrea Guzman, Shamak Gowda, Samhita Konduri, Jiachen Lian, Robin Netzorg, Gopala Anumanchipalli
Audio & Speech
  • TART addresses limitations in existing guitar transcription systems by capturing expressive techniques.
  • The modular design allows for improved accuracy in string-fret assignments and generalization to noisy audio.
  • TART achieves state-of-the-art results in zero-shot evaluations across multiple benchmarks.
  • The framework generates tablature with detailed fingering and expressive annotations directly from audio.
Read more
When More Is Not Better: Component Anti-Synergy in a P300 Speller
Lucas Yang, Rui Liu, Fusheng Wang
Theory
  • Component contributions in P300 spellers are conditional rather than additive.
  • Calibration is the most crucial component for performance in P300 spellers.
  • Adding components can lead to anti-synergy, reducing overall system performance.
  • Language model support is not universally beneficial and depends on EEG evidence quality.
Read more
Beyond Solver Verdicts: Generative Reward Models for Autoformalization
Vikash Singh, Debargha Ganguly, Aman Goel, Ali Torkamani, Xiaoxue Han, Joseph Lilien, Ferhat Erata, Vipin Chaudhary
Theory Generative Models NLP
  • Introduction of Verdict-Preserving-Unfaithfulness (VPU) as a formalized failure mode in autoformalization.
  • Development of Generative Verification (GenV) that distills offline equivalence checks into a continuous scoring system.
  • Empirical validation showing GenV's effectiveness in detecting VPU across diverse formal styles and unseen translators.
  • Achieved an AUROC score of 0.961 in reference-equivalence verification, demonstrating high accuracy.
Read more
CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
Theory Graph Learning
  • CausalArena introduces a unified benchmark for causal discovery, addressing inconsistencies in evaluation protocols.
  • The benchmark includes synthetic, semantic operational, and formula-grounded SCMs to provide diverse evaluation environments.
  • Experiments show substantial ranking shifts among methods across different SCM families, indicating the complexity of causal discovery evaluation.
  • CausalArena supports the introduction of new causal environments as pretrained models evolve, ensuring relevance and adaptability.
Read more
When does a spectral prior help graph learning? Connectivity-loss estimation under road-network disruptions
Van-Truong Le
Graph Learning
  • Introduces a hybrid estimator for connectivity loss in road networks using GNNs.
  • Demonstrates the effectiveness of a spectral prior in improving graph learning under specific conditions.
  • Finds that the benefits of the spectral prior are context-dependent and not universally superior.
  • Highlights the limitations of using a residual prior in the presence of available training data.
Read more
A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph
Ruben Cartuyvels, Karim Douch, Gabriele Bertoli, Mounia El Baz, Artemis Vrettou, Sรฉbastien Lefรจvre, Diego Fernandez Prieto
Graph Learning Time Series Optimization
  • Introduction of AmazonWSE, a dataset for WSE imputation covering 19K river sections in the Amazon basin.
  • The dataset presents extreme sparsity, with less than 1% of sections observed daily, challenging existing imputation methods.
  • A novel bidirectional selective state space model is proposed, outperforming traditional methods by leveraging subgraph sampling.
  • The model achieves significant RMSE reductions compared to state-of-the-art methods, providing broader coverage and accuracy.
Read more
DR-LabStack: Design and Implementation of a Clinician-Facing Web System for Diabetic Retinopathy Prediction
Yingfan Xu, Tieming Liu, Ye Liang
Computer Vision
  • DR-LabStack integrates multiple pretrained diabetic retinopathy prediction models into a single web interface.
  • The system addresses the challenges of heterogeneous input requirements and preprocessing for different models.
  • Functional evaluations confirmed the successful integration and operational consistency of the models.
  • The design emphasizes a clinician-friendly interface while maintaining model-specific distinctions.
Read more
Halo: Improving forecast accuracy through heteroscedastic estimation
Adam Cataldo
Time Series
  • Halo enhances point estimate accuracy in time series forecasting through heteroscedastic estimation.
  • The method involves modifying existing deep learning architectures to include scale parameter estimation.
  • Significant improvements in forecast accuracy were observed across multiple models and metrics.
  • The approach is effective even with hyperparameters tuned for baseline point estimates.
Read more
RDDMPI: Residual Denoising Diffusion Model for Probabilistic Multivariate Time Series Imputation
Ramiro Valdes Jara, David Chapman, Adam Meyers
Time Series
  • RDDMPI reformulates MTSI as a baseline-residual decomposition, separating deterministic reconstruction from probabilistic uncertainty modeling.
  • The framework operates in residual space, allowing for focused correction of systematic errors.
  • A reliability-aware conditioning mechanism is introduced to manage the influence of baseline predictions adaptively.
  • RDDMPI shows improved performance over existing methods in reconstruction accuracy and uncertainty estimation.
Read more
Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
Jae Gon Kim, Donghoon Yoo, Hanyul Ryu, Sungho Ha, Juyeon Lee, Soojung Ryu
Large Language Models Optimization Efficient ML
  • Optimal power settings depend on the specific model and hardware combination, not just GPU class.
  • Disaggregated architectures allow for distinct power profiles for prefill and decode phases.
  • The proposed model-calibrated controller outperforms traditional static profiles in efficiency and latency management.
  • Dynamic calibration can convert latency headroom into energy savings effectively.
Read more
Flow Duality and Source Geometry for Categorical Generation
Etrit Haxholli
Generative Models Theory NLP
  • Establishes a duality between continuous and discrete flow matching for categorical generation.
  • Demonstrates how different source geometries affect transition timing and vocabulary size dependence.
  • Derives discrete interpolation behavior for Gaussian, bounded-uniform, and centered negative-exponential sources.
  • Highlights the significance of continuous source distribution as a design choice in generative modeling.
Read more
Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study
Yan Hon Michael Chung, Hanlin Wang
Computer Vision NLP Multimodal
  • Combining synthetic and real data significantly improves OCR accuracy for low-resource languages.
  • Real training images raise the performance of VLMs to 95.09-96.28% accuracy, compared to 87.92% for synthetic-only models.
  • Joint and sequential training methods yield similar results in archival accuracy.
  • A compact CRNN achieves high performance when real images are included, indicating model scale is not the sole determinant of accuracy.
Read more
Conformal Calibration Transfer
Achref Doula
Theory Multimodal Efficient ML
  • Introduces Transported Conformal Calibration (TCC) for transferring calibration from source to target space.
  • Utilizes unlabeled paired observations to transport labeled calibration effectively.
  • Offers two correction methods: TCC-KS for conservative adjustments and weighted-TCC for efficiency.
  • Demonstrates reliable target-domain coverage transfer across multiple datasets without labeled target data.
Read more
Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models
Ken Chen, Maneesha Perera, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge
Multimodal Time Series Large Language Models
  • Introduction of SolCloudLLM, a multimodal framework for solar forecasting.
  • Effective integration of sky images and time-series data enhances forecasting accuracy.
  • Achieved a maximum relative MSE reduction of 25.4% on the SIRTA dataset.
  • Demonstrated superior performance in few-shot learning scenarios.
Read more
Relatively Smart II: Tractable or Semi-Supervised Instance-Optimal Learning
Shaddin Dughmi, Alireza F. Pour
Theory
  • ERM and any proper consistent learner are shown to be relatively smart for binary classification.
  • A semi-supervised relatively smart learner can achieve quadratic blowup only in unlabeled sample complexity.
  • The proposed learner generalizes the One-Inclusion Graph to a leave-most-out transductive problem.
  • The trade-off between efficiency and tractability is highlighted, with implications for oracle calls in semi-supervised learning.
Read more
MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions
Antoine Saillenfest
NLP Large Language Models Interpretability
  • Introduction of MUtE*, a class of optimal erasure functions that defines a dual counterfactual mapping.
  • Development of a computationally efficient implementation that utilizes a translational bias for erasure and counterfactual generation.
  • Empirical validation shows significant improvements in bias mitigation and counterfactual text generation.
Read more
Phases in a class of associative memories via hidden neurons
Toshihiro Ota, Masato Taki
Theory
  • Introduces a bipartite architecture for associative memory that integrates hidden neurons as an order parameter.
  • Analyzes retrieval dynamics under polynomial and exponential loads using statistical mechanics methods.
  • Identifies distinct retrieval phases and their statistical behaviors based on the type of visible neurons.
  • Demonstrates the role of Lagrangians in fixing stability and storage capacity in associative memory models.
Read more
GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models
Xuan Cuong Ngo, Hao Vo, Ngan Le
NLP Large Language Models Optimization
  • GEOSTEER formulates activation steering as a Riemannian optimization problem.
  • It replaces fixed one-step updates with adaptive, multi-step geodesic updates.
  • The method improves stability and expressiveness of steering while preserving activation norms.
  • Experiments show consistent performance improvements over existing activation steering baselines.
Read more
Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance
Daichi Kuroda, Maximilien Dreveton, Matthias Grossglauser, Patrick Thiran
Theory
  • Hierarchical clustering can satisfy Kleinberg's axioms of scale invariance, richness, and consistency, unlike flat clustering methods.
  • The authors construct several admissible hierarchical clustering methods, demonstrating the existence of uncountably many such methods.
  • A partial order exists among admissible methods, revealing substantial diversity and the absence of a greatest method.
  • Every admissible method contains a common backbone of sufficiently well-separated clusters.
Read more
Counterfactual Marginalisation: Framework for Evaluating Robustness to Nuisance Variables
Yasin Ibrahim, Hermione Warr, Robin J. Evans, Konstantinos Kamnitsas
Generative Models Computer Vision Interpretability
  • Introduction of counterfactual marginalisation as a test-time evaluation procedure for medical image classifiers.
  • Development of new evaluation metrics that assess model robustness to nuisance variables without requiring disease labels.
  • Demonstration of the framework's ability to expose biases in predictions better than traditional metrics.
  • Emphasis on the importance of evaluating model sensitivity, stability, and worst-case performance in clinical settings.
Read more
Reification as a Transferable Vocabulary: Zero-Shot Link Prediction with Vanilla GNNs
Camille Pradel
Graph Learning
  • Reification transforms facts into nodes, enabling a shared vocabulary across datasets.
  • Five standard GNNs can achieve zero-shot transfer to multiple benchmarks using the proposed representation.
  • The GAT model matches the performance of ULTRA, a specialized foundation model, in several evaluations.
  • The approach is applicable to relational databases, showcasing its broad utility.
Read more