AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

48 Papers today
8h Update frequency
7 Days of history
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia
Large Language Models Reinforcement Learning Optimization
  • CalibForge synthesizes terminal tasks using adversarial solver calibration to ensure tasks are appropriately challenging.
  • Two calibration strategies (multi-solver and contrastive) enhance the task validation process by focusing on solver behavior.
  • The system generated 5,431 calibrated tasks, significantly improving model performance on multiple benchmarks.
  • Models trained on calibrated tasks outperformed baseline models by substantial margins, demonstrating the effectiveness of the approach.
Read more
Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated Learning
Alireza Moayedikia, Alicia Troncoso Lora
Federated Learning
  • Client heterogeneity estimation from updates is confounded by device capacity.
  • Adaptive capacity allocation can lead to model corruption if clients are under-resourced.
  • A coverage guarantee can prevent failure modes in adaptive allocation strategies.
  • Uniform allocation may perform comparably to heterogeneity-aware policies under certain conditions.
Read more
ArchEGraph: A Large-Scale Graph Dataset for Geometry-Topology-Physics Aligned Building Energy Modeling
Yihui Li, Yihui Chen, Kaidi Zha, Xiaoyue Yan, Zhexuan Yu, Shiqi Dai, Jun Xiao, Jun Yin, Ramon Elias Weber, Borong Lin
Graph Learning
  • ArchEGraph is a large-scale dataset that aligns building geometry, topology, and physics for energy modeling.
  • The dataset includes 5,481 buildings and over 49,000 simulation cases, capturing complex geometric and topological features.
  • Two benchmark tasks are defined: graph reconstruction from meshes and topology-informed load prediction.
  • Standardized evaluation protocols are introduced to assess model performance across various conditions.
Read more
Faster Query-Key Learning Sharpens Attention in Self-Attention Models
Rahul Vashisht, Harish G. Ramaswamy
NLP Theory Interpretability
  • The parameterization of query-key and output-value circuits significantly affects attention patterns in self-attention models.
  • Faster learning rates for query-key parameters lead to sharper attention on task-relevant tokens.
  • The study provides closed-form dynamics that explain how relative learning speeds shape attention during training.
  • Experiments confirm that attention structures can be altered without compromising predictive performance.
Read more
CohortHijack: Robustness of Single Cell Annotation to Companion Cell Removal
Arash Vashagh, Yasmin Vashagh
Theory
  • CohortHijack identifies companion-cell removal as a threat to single-cell annotation integrity.
  • Structured removal methods consistently outperform random removal strategies.
  • Multi-start search techniques can significantly alter target annotations with minimal collateral impact.
  • Cohort composition is a critical factor in the reliability of single-cell annotations.
Read more
Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs
Hamed Damirchi, Ignacio Meza De la Jara, Damith Ranasinghe, Yuhang Liu, Javen Shi
NLP Large Language Models Interpretability
  • Introduces a three-stream detector for improved reasoning error detection in LLMs.
  • Combines motion with coarse and fine location readings to enhance context interpretation.
  • Achieves up to 12% accuracy improvement over existing displacement-only methods.
  • Demonstrates effectiveness across various reasoning benchmarks, including factual tasks.
Read more
Sub-Quadratic Bisimulation Metrics via Approximate Nearest Neighbors: Coverage-Augmented Guarantees and Computable Two-Sided Certificates
Ibne Farabi Shihab, Joyanta Jyoti Mondal
Reinforcement Learning Theory Efficient ML
  • Introduces a sub-quadratic algorithm for computing bisimulation metrics in MDPs.
  • Utilizes an approximate nearest neighbor index for efficient state pair selection.
  • Establishes a coverage-augmented anytime error bound for the proposed method.
  • Demonstrates significant computational efficiency improvements in experiments.
Read more
How Far Do Simple Transformations Translate Across Text Embedding Models?
Sid Ali Hamideche, Louis Adrien Dufrene, Quentin Lampin, Guillaume Larue
NLP Theory Interpretability
  • Simple transformations can recover shared structures in some text embedding models but not universally.
  • Compatibility of transformations is influenced by model architecture, training objectives, and pooling strategies.
  • The study employs a multi-diagnostic approach, evaluating geometric similarity, retrieval, and downstream transfer.
  • Findings challenge the literature's assumption of universal latent compatibility across different models.
Read more
Mathematical Principles and Experimental Discoveries of the Emergence of Symbolic Patterns in Artificial Neural Networks
Quanshi Zhang, Qihan Ren, Siyu Lou
Theory Interpretability
  • Symbolic patterns can emerge in ANNs, providing a means to explain their inference logic.
  • Two mathematical criteriaβ€”monotonicity and smoothnessβ€”lead to the emergence of sparse symbolic interactions.
  • The emergent symbolic patterns demonstrate strong transferability across different input samples and models.
  • The study suggests a shift towards communicative learning, allowing for direct inspection and tuning of ANN logic.
Read more
Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines
Dohyeon Kong, Jaebong Cho, Hyunbo Cho
Computer Vision Robotics
  • Introduction of an equipment-centric localization framework for hot forging environments.
  • Utilization of video streams and event-driven FSMs for robust workpiece tracking.
  • Achieved 100% event detection accuracy and a mean localization error of 317.8 mm.
  • Integration of KPGA mechanism enhances activity recognition performance.
Read more
Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics
Hailong Jiang, Emran Hossain, Feng Yu, Jianfeng Zhu, Guilin Zhang, Wulan Guo
Theory Reinforcement Learning
  • Introduction of Quantum-Structured World Models (QSWMs) as a quantum-inspired framework for predictive modeling.
  • Establishment of three foundational properties: classical inclusion, predictive sufficiency, and structured compactness.
  • Demonstration of the effectiveness of ComplexQSWM over classical baselines in local predictive tasks.
  • Identification of limitations in long-horizon predictions and latent interpretability.
Read more
MAUPITI: On-Device Prototype-Based Learning on a Smart Infrared Sensor
Beatrice Alessandra Motetti, Tanguy Dugas du Villard, Matteo Risso, Alessio Burrello, Francesco Daghero, Enrico Macii, Massimo Poncino, Marco Castellano, Alfio Basile, Daniele Jahier Pagliari
Computer Vision Efficient ML
  • MAUPITI integrates a low-power infrared sensor with on-device learning capabilities.
  • The use of a prototype-based NCM classifier allows for efficient online adaptation without backpropagation.
  • The system operates within tight memory and power constraints, making it suitable for embedded applications.
  • Experimental results show comparable accuracy to traditional classifiers with minimal latency.
Read more
Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
Moloud Damandeh, Meead Saberi
Multimodal
  • Introduces a new dataset linking sidewalk-view imagery with individual walkability ratings.
  • Demonstrates that the viewpoint of imagery (sidewalk vs. street view) significantly affects walkability ratings.
  • Proposes a multimodal deep learning framework that incorporates user attributes to model subjective variability in walkability perception.
  • Achieves a 65% improvement in rank agreement over traditional image-only models.
Read more
SkillTFM: Gated Skill Evolution for Training-Free Adaptation of Tabular Foundation Models
Yi He, Zhengkang Guan, Anpeng Wu, Peng Cui, Fei Wu, Kun Kuang
Time Series Optimization Efficient ML
  • SkillTFM is the first skill-based adaptation system for training-free tabular foundation models.
  • It employs a gated skill evolution mechanism that couples selective repairs with safe fallbacks.
  • The system demonstrates significant improvements in prediction accuracy, particularly in boundary scenarios.
  • SkillTFM's learned skill state is transferable across different TFM backbones and optimizer settings.
Read more
Align-RAG: Alignment Is All You Need for TSFM In-Context Learning
Mohammad Asadi, Soheil Hor, Bardiya Akhbari, Jack W. O'Sullivan, Tahoura Nedaee, Layne C. Price, Raviteja Anantha, Euan Ashley, Ehsan Adeli
Time Series
  • Align-RAG is a training-free method that enhances frozen TSFMs for time series forecasting.
  • It outperforms the state-of-the-art TS-RAG method across multiple datasets without requiring learned parameters.
  • The method applies amplitude rescaling and phase shifting to align retrieved data with the query.
  • Align-RAG improves zero-shot forecasting accuracy significantly across various TSFM architectures.
Read more
A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies
Frieder Wizgall, Georg Tirpitz, Moritz Seiler, Kerstin Ritter, BΓ‘lint MucsΓ‘nyi
Theory Interpretability
  • Introduces a unified definition of uncertainty as pointwise posterior risk.
  • Develops a benchmark for direct computation of oracle epistemic and aleatoric uncertainty.
  • Demonstrates that accurate predictions do not guarantee reliable uncertainty estimates.
  • Highlights the limitations of existing proxy evaluations for uncertainty assessment.
Read more
SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models
Hoda Fakharzadehjahromy, Emil Wiman, Andreas Bueff, Hafsteinn Einarsson, Fredrik Heintz
NLP Large Language Models Reinforcement Learning
  • SAGA utilizes dependency-parser supervision to replace costly human preference annotations.
  • The framework improves grammatical quality in low-resource languages without requiring human labels.
  • Parser-derived supervision effectively addresses challenges such as reward hacking and alignment tax.
  • Significant improvements in grammatical accuracy were observed across Danish, Icelandic, and Norwegian BokmΓ₯l.
Read more
The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity
Iosif Lytras, Nikolaos Makras, Sotirios Sabanis
Optimization Theory Large Language Models
  • Introduction of SG-TULA, a novel algorithm for sampling from non-convex distributions with non-smooth potentials.
  • Derivation of non-asymptotic convergence bounds in Wasserstein-2 distance with explicit constants.
  • Demonstration of SG-TULA's effectiveness in pretraining LLMs, achieving competitive results against established optimization methods.
  • Addressing the challenges of superlinear gradient growth and non-convexity in optimization problems.
Read more
Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration
Ruochen Jin, Zhanliang Wang, Zongyu Dai, Jiancong Xiao, Bojian Hou
NLP Large Language Models Optimization
  • Identifies the limitations of post-hoc temperature scaling in LLM calibration.
  • Proposes a bilevel optimization framework for training-time calibration adjustments.
  • Introduces an entropy-maximization objective to mitigate overconfidence in LLMs.
  • Demonstrates strong calibration performance across various tasks and datasets.
Read more
When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters
Fangxin Wang, Ziyi Zhang, Diyi Zhuang, Langzhou He, Shiyu Wang, Baichuan Mo, Philip S. Yu
Time Series Interpretability Large Language Models
  • CRAFTER introduces a source-blind framework for corrective feature discovery, focusing on model-failure processes.
  • The framework combines two feature generators: a compositional search and a large language model.
  • CRAFTER significantly outperforms existing feature-engineering systems across multiple datasets and models.
  • The effectiveness of corrective features is regime-dependent, highlighting the need for careful evaluation.
Read more
A Rate Separation for Agnostic Direct Sums
Mihir More, Aritra Das, Debayan Gupta
Theory
  • The single-instance learning rate does not determine the direct-sum learning rate.
  • Agnostic learning curves for constant functions and zero/identity functions are of order n^(-1/2).
  • The paper provides a rate separation theorem demonstrating differing rates for direct sums as the number of factors increases.
  • The results challenge existing assumptions about the relationship between single-instance and direct-sum learning rates.
Read more
Accelerating nanodrug development in continuous flow systems using informed prediction models based on low-cost surrogate nanoparticles
Kai Dahms, Eilien Heinrich, Jochen Schmid, Michael Bortz, Iryna Savych, Regina Bleul
Optimization
  • Introduction of a predictive modeling approach based on shape constraints for nanoparticle development.
  • Utilization of controlled microfluidic methods to systematically prepare liposomes and lipid nanoparticles.
  • Validation of the model with minimal empirical data, showcasing its effectiveness in predicting nanoparticle characteristics.
  • Reduction of experimental workflows, leading to cost and time efficiency in nanodrug development.
Read more
Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction
Gregor Molan, Grafika Jati, Francesco Barchi, Andrea Acquaviva, AljaΕΎ Osterman, Martin Molan
Time Series Optimization Efficient ML
  • Introduction of the FSD-RM paradigm for small-data representation learning.
  • Use of dimension-aware neural architecture search to optimize model design.
  • Demonstration of competitive predictive performance on cryocooler telemetry data.
  • Focus on capacity-controlled representation learning without large-scale pretraining.
Read more
BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells
Yuhao Wang, Zelin Zang, Yuxuan Liu, Zhen Lei, Stan Z. Li
Graph Learning
  • BioM-JEPA predicts aggregate representations of graph-connected gene blocks instead of individual genes.
  • The model employs a student-teacher framework to enhance representation learning efficiency.
  • Linear attention is used to manage gene interactions, improving computational efficiency.
  • BioM-JEPA outperforms existing models in retaining biological information and reducing errors in perturbation-response tasks.
Read more
Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization
Grigory Malinovsky
Optimization Federated Learning Theory
  • Introduction of ProxSkip for communication-efficient local gradient updates.
  • Development of Variance Reduced ProxSkip to balance communication and computation costs.
  • Demonstration of scalability under partial client participation.
  • Establishment of a mechanism for Byzantine robustness through gradient difference clipping.
Read more
Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning
Wei-Hsiang Chen, Pin-Hsuan Yu, Chen-Hsuan Fang, Jung-Hua Wang
Computer Vision Theory Efficient ML
  • Identification of granularity bias as a factor influencing removal budgets in adaptive data cleaning.
  • Development of a matched operating-point evaluation framework to separate true detection capability from removal count effects.
  • Demonstration that naive evaluations can misattribute performance improvements to better corruption discrimination.
  • Findings indicate that performance gaps largely disappear when operating points are equalized, especially at low-to-moderate corruption rates.
Read more
Beyond Attention: Signed Integrated Gradients Attribution in a BiomeGPT-Style Microbiome Transformer
Oren Nelson
Interpretability
  • Introduces signed Integrated Gradients as a solution to the limitations of [CLS] attention weights in BiomeGPT.
  • Establishes a fusion-aware baseline that preserves species identity while isolating abundance effects.
  • Demonstrates how the proposed method reveals directional relationships between microbial species and health outcomes.
  • Recommends the use of second-order Integrated Hessians for understanding interactions among microbiome members.
Read more
Hidden Gauge Controls Feature Specialization in ReLU Networks
Tongxi Wang
Theory Optimization
  • Demonstrates a Θ(DΒ²) separation in specialization times for functionally identical neurons under different hidden gauges.
  • Establishes that a favorable gauge can deterministically assign feature ownership to one neuron while rendering others redundant.
  • Introduces a mechanism that separates changes in functional coefficients from changes in feature direction.
  • Validates results through population and finite-sample training, showing robustness to visible perturbations.
Read more
EpiFlow: A framework for improving the utility of wastewater signals for disease forecasting
Aniruddha Adiga, Jingyuan Chou, Gursharn Kaur, Andrew Warren, Srinivasan Venkatramanan, Baltazar Espinoza, Bryan Lewis, Justin Crow, Alexandra Lorentz, Rekha Singh, Madhav Marathe
Time Series
  • EpiFlow improves the forecasting accuracy of disease burden indicators using wastewater viral load signals.
  • The framework incorporates advanced data preprocessing and signal analysis techniques to handle noise and variability in wastewater data.
  • Forecasting models demonstrate a 20 percentage point improvement in forecast coverage during critical phases.
  • The utility of wastewater signals is maintained even during low-prevalence periods and reporting delays.
Read more
Surv-IPTB: An Attention-Based Model for Estimating Individual Probability of Treatment Benefit with Survival Data
Lev V. Utkin, Stanislav K. Kogan, Andrei V. Konstantinov
Theory
  • Surv-IPTB reformulates IPTB estimation as a binary classification problem, enhancing individual treatment benefit assessments.
  • The model incorporates an attention mechanism to effectively aggregate pairwise patient comparisons and handle censored data.
  • Extensive experiments show superior performance of Surv-IPTB over traditional meta-learner baselines in complex nonlinear scenarios.
  • The approach provides a principled method for estimating treatment benefits tailored to individual patients, addressing limitations of average treatment effect assessments.
Read more
IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games
Conor M. Artman, Nicholas Di, Scott Perkins
Reinforcement Learning Generative Models Theory
  • IFlowNets generalize AFlowNets to handle incomplete information games effectively.
  • The paper proves that existing generative flow network constraints are inadequate for incomplete information settings.
  • IFlowNets maintain essential properties like flow matching, crucial for achieving valid player strategies.
  • Preliminary results show IFlowNets outperform or match the performance of established methods in standard game environments.
Read more
PPDL: LLM-Based Flows as Probabilistic Programs
Louis Mandel, Guillaume Baudart, Mandana Vaziri, Martin Hirzel
Large Language Models NLP Theory
  • Introduction of PPDL, the first probabilistic programming language for LLM-based flows.
  • Decoupling of inference scaling from core program logic, enhancing flexibility and usability.
  • Formal semantics that clarify the interaction between prompt-based sampling and probabilistic factors.
  • Empirical results demonstrating the effectiveness of PPDL with various inference engines.
Read more
QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding
Ayushman Garg, Akshita Gupta, Shaswata Bhattacharya, Abhishek Gupta, Sandeep Kumar, Manoj Kumar
NLP Large Language Models Efficient ML
  • QEvict introduces a recoverable eviction strategy that allows for dynamic management of KV cache, addressing the limitations of traditional irreversible eviction methods.
  • The method categorizes token windows into three tiers, enabling the retention of important contexts while maintaining a fixed memory budget.
  • QEvict effectively reduces missed attention and improves information retention in long-context decoding tasks.
  • The proposed diagnostics, Future Missed Mass and Global LIR, provide insights into the importance of cached states over time.
Read more
When Does Consensus Mean Correctness? Measuring the Agreement-Accuracy Coupling with Semantics-Preserving Re-Rendering
Rasul Khanbayov, Hasan Kurban
Computer Vision Large Language Models Theory
  • Introduces RENDEQ, a tool for generating semantically equivalent renderings for accurate measurement of model correctness.
  • Demonstrates that re-rendering is superior to resampling in assessing model accuracy and reliability.
  • Finds that model agreement does not always correlate with correctness, especially when errors are diffuse.
  • Identifies that fine-tuning on consensus can lead to decreased accuracy, contrary to existing literature.
Read more
Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers
Zhen Zhang, Amr Alanwar
Theory Efficient ML Optimization
  • Introduces Matrix Zonotopic Attention (MZAttn) for improved set transformer performance.
  • Defines Transformation Degrees of Freedom (TDOF) to analyze the complexity of target operators.
  • Demonstrates that MZAttn can represent complex targets with fewer layers compared to traditional attention mechanisms.
  • Experimental results indicate significant performance improvements on high-complexity tasks.
Read more
Online Security Learning in Cooperative Multi-Agent Systems under Hidden Byzantine Attacks
Ximing Sun, Yue Wang
Theory Reinforcement Learning Robotics
  • Introduces a Byzantine team MDP framework to model multi-agent systems under hidden Byzantine attacks.
  • Establishes the concept of security regret, decomposing it into return regret and response gap.
  • Demonstrates that public feedback cannot certify security against unrestricted attacks.
  • Develops a robust estimation-to-decisions learner with a proven regret bound.
Read more
Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors
Saeid Shokoufa, Mohammad Erfan Sadeghi, Mehdi Kamal, Massoud Pedram
Large Language Models Efficient ML Optimization
  • Introduction of HYMELL, a hybrid framework for estimating LLM inference latency and energy.
  • Three-level modeling approach combining analytical and machine learning techniques.
  • High predictive accuracy achieved, with less than 5% error for LLaMA 3 8B model.
  • Framework supports diverse architectures, enhancing its applicability.
Read more
CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions
Shuheng Cao, Zhenhao Zhang, Ruiqi Chen, Renjie Cao, Weijia Zhang, Siyu Zhang, Jiaxin Liu, Xiangyu Zeng, Haotian Geng, Fan Gu
Multimodal Theory Graph Learning
  • CertBind enables certifiable composition of multimodal connectors, enhancing decision-making capabilities.
  • The framework operates at multiple scales, addressing ambiguities in retrieval outputs.
  • Different deployment actions are defined based on the support of retrieval routes.
  • The methodology includes contract-aware conformal ranks and a majority path-transversal budget for robustness.
Read more
When GNNs Fail: Quantifying and Overcoming Temporal Correlation Volatility in Time Series
Chen Shao, Yue Wang, Zhenyi Zhu, Zhanbo Huang, Tobias KΓ€fer, Zonghan Wu, Danai Koutra
Graph Learning Time Series
  • Introduces Temporal Correlation Volatility (TCV) as a metric for quantifying instability in time series correlations.
  • Demonstrates that existing GNN models struggle in high-TCV environments, leading to significant performance degradation.
  • Proposes GLIDE, a novel GNN layer that effectively addresses dynamic graph structures through innovative design mechanisms.
  • Shows that GLIDE outperforms existing models by up to 45.6% on average, with gains reaching 85.7% in specific scenarios.
Read more
Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction
Quinn Ledingham, Zhengsen Xu, Yimin Zhu, Zack Dewis, Mabel Heffring, Saeid Taleghanidoozdoozan, Motasem Alkayid, Megan Greenwood, Lincoln Linlin Xu
Interpretability
  • TabPFN outperforms other models in predicting post-wildfire debris flows.
  • Short-duration rainfall intensity and storm accumulation are the most important features for prediction.
  • Synthetic data augmentation significantly improves model performance.
  • The study provides a systematic evaluation framework for machine learning models in hazard prediction.
Read more
Quantization Damage Is Multiplicative, Not Additive
Zekun Wu, Swati Dhiman, Adriano Koshiyama
NLP Large Language Models Theory
  • Quantization damage is characterized as a multiplicative loss of decision margin rather than additive noise.
  • The concept of 'margin shrinkage' explains how quantization affects model decisions, particularly in safety-critical applications.
  • A predictive model for decision flip probabilities is developed, showing high accuracy without relying on fitted flip data.
  • The findings challenge existing strategies that focus on protecting specific weights during quantization.
Read more
Threshold-Based Early Stopping of Accumulations in Neural Networks with Binary Activation
Quentin Luquet de Saint-Germain, Massil Ait Abdeslam, Jean Pierre David
Efficient ML
  • Introduces a post-training early-stopping mechanism for binary neural networks.
  • Demonstrates significant reductions in computational operations without retraining model parameters.
  • Achieves 86.6% reduction in accumulation terms with minimal accuracy drop.
  • Focuses on the efficiency of AI in constrained environments, addressing both accuracy and resource usage.
Read more
Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression
Haolin Tian, Yuzhe Liu, Tonghan Wang
Large Language Models Optimization Efficient ML
  • Introduces GraceKV, a global approach for KV cache compression that optimally allocates resources across layers and context slots.
  • Utilizes a prototype tree structure to represent KV entries, allowing for adaptive balancing of resolution and coverage.
  • Demonstrates superior performance compared to existing methods, achieving high compression ratios without additional training.
  • Validates the approach through systematic experiments across diverse long-context tasks, showing effectiveness in resource allocation.
Read more
EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
Xuying Ning, Dongqi Fu, Tianxin Wei, Hanqing Zeng, Yuanchen Bei, Bingxuan Li, Zihao Li, Qifan Wang, Xiang Shen, Yifan Wu, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He
Large Language Models Reinforcement Learning Robotics
  • EvoHarness-RL introduces a self-evolving runtime harness for long-horizon LLM agents.
  • The framework abstracts harness components into a unified Belief, Progress, and Experience (BPE) state.
  • A two-stage training process enhances the agent's ability to construct and utilize external state effectively.
  • The approach significantly improves task success rates and efficiency in long-horizon interactions.
Read more
How Molecular Generative Models Organize Molecular Identity
Raul Ortega-Ochoa, Tejs Vegge, Jens S. Bakander, Luis Mantilla Calderon, Alan Aspuru-Guzik, Tonio Buonassisi
Generative Models Graph Learning Theory
  • Molecular generative models exhibit a structured internal organization of chemical identities.
  • The arrangement of identities is influenced by representation type, identity conventions, and stochasticity in decoding.
  • Local chemical cohesiveness stabilizes during training, but the number of distinct identities in neighborhoods changes.
  • Understanding the internal organization is crucial for treating generative spaces as chemically navigable.
Read more
Fluid-DiT: Graph-Free Diffusion Transformers for Fluid Flow Simulations Learning
Shentong Mo, Guolin Ke
Generative Models Graph Learning Efficient ML
  • Fluid-DiT is the first graph-free diffusion transformer for fluid flow simulations, enhancing local precision and long-range coupling.
  • The model employs a latent-space diffusion formulation that improves training efficiency and reduces high-frequency artifacts.
  • Fluid-DiT consistently outperforms state-of-the-art graph-based models in distributional fidelity and scalability.
  • The framework generalizes effectively across diverse geometries and Reynolds numbers without requiring mesh-specific tuning.
Read more
Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping
Vaishnav Vaidheeswaran, Dilith Jayakody, Biruk Ambaw, Jaswanth Kumar, Md Mahbub Alam, Gabriel Spadon
Reinforcement Learning Robotics Theory
  • Latent context in IRL does not necessarily improve performance and may reduce it in certain scenarios.
  • Observable route and environmental conditions explain most behavioral variations in Arctic shipping.
  • Nonlinear reward models significantly outperform linear models in predicting vessel behavior.
  • A context-need diagnostic is proposed to evaluate the necessity of latent context in decision-making.
Read more
SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction
Guofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu, Amelie Chi Zhou
Graph Learning Efficient ML Theory
  • Introduction of SNI-GNN, a SmartNIC-assisted system for full-graph GNN training.
  • Reduction of inter-node communication by 21-45% through in-network embedding prediction.
  • Achieved end-to-end speedups of 1.3-3.6x over existing full-graph training systems.
  • Maintained accuracy loss within 0.01 while scaling to 16 GPUs on large graphs.
Read more