Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
Token pruning framework for efficient video vision-language models reducing computational cost via temporal token scoring.
Token pruning framework for efficient video vision-language models reducing computational cost via temporal token scoring.
Graph neural network method for air quality forecasting modeling pollution diffusion between cities and monitoring stations.
Aergia: federated learning system leveraging client heterogeneity in computing power to reduce training time.
Feature space renormalization mechanism for semi-supervised learning improving consistency regularization on unlabeled data.
Soft Dice Confidence: confidence estimator for selective prediction in semantic segmentation enabling model abstention.
Hi-GMAE: hierarchical graph masked autoencoders for multi-scale self-supervised learning on graph-structured data.
Analyzes transformer-based amortized causal discovery on observational data, bridging supervised learning with identifiability theory.
Den-TP: data curation framework addressing long-tail distribution in trajectory prediction datasets for autonomous driving.
ACT-JEPA: joint-embedding predictive architecture for efficient policy representation learning via self-supervised learning from unlabeled data.
SALSA-RL: stability analysis method for deep reinforcement learning agents enabling interpretability and safety assessments in continuous control.
Studies regret minimization in repeated first-price auctions with causal inference for online advertising scenarios.
Offline reinforcement learning algorithm leveraging inverse optimization and sub-optimality loss for continuous state/action spaces.
Hierarchical federated learning framework using UAVs as mobile aggregators for distributed IoT systems with limited connectivity.
Proposes minimal repair concept showing imputing all missing values unnecessary; identifies critical missing data subsets for accurate ML models.
SocialJax: evaluation suite for multi-agent reinforcement learning in sequential social dilemmas, measuring agent generalization.
Proposes Arch-VQ for learning discrete neural architecture representations using autoregressive priors instead of continuous VAE mapping.
Studies impact of duplicated training data on deep neural network image classifiers, comparing robust vs. standard models against adversarial attacks.
arXiv paper: ILLUME method for post-hoc explainability of tabular ML models with interpretable meta-encoding.
arXiv paper: Clust-Splitter algorithm for efficient clustering on large datasets using nonsmooth optimization.
arXiv paper: Bi-level policy optimization with Nyström hypergradients for actor-critic reinforcement learning algorithms.
arXiv paper: Restoration Score Distillation framework for learning generative models from corrupted data. Novel ML research.
Foundation model for time series forecasting using mixture-of-experts architecture with decoupled training to handle diverse temporal patterns and multi-variable correlations.
Tutorial on diffusion and flow-based generative models covering mathematical foundations, ODEs, SDEs, and core algorithms for image, video, and multi-modal generation.
Offline reinforcement learning method for mismatched dynamics leveraging model-based approaches to explore high-reward states.
Learning to Reject framework extending ML models to abstain from predictions and explanations with low quality.
Fast weight programmers with 2D matrix hidden states connecting RNNs, language modeling, and neurobiology.
Tree-based group relative policy optimization for LLM agents addressing sparse supervision in multi-turn tasks.
Method aligning supervised fine-tuning with in-context learning activations to improve LLM generalization and calibration.
Offline reinforcement learning using in-context learning with linear Transformers for compositional Q-function estimation.
Parameter-efficient unlearning method for foundation models addressing privacy/safety with bounded weight growth.
Slow-Fast Policy Optimization framework for improving LLM reasoning via reinforcement learning with stable gradient updates.
Proposes continual low-rank adapters for LLM-based recommender systems handling evolving users and preferences without catastrophic forgetting.
Presents augmentation-free graph contrastive learning via fractional-order neural diffusion networks for multi-scale structure learning.
Proposes standardized methodology for evaluating long-term sustainability and efficiency of ML models addressing Green AI gaps.
Demonstrates in-context learning emerges organically in genomic sequence models trained with next-token prediction on DNA sequences.
Develops methods for provably safe ML model updates preventing catastrophic forgetting and alignment drift in dynamic environments.
Proposes data filtering method for cross-domain offline RL addressing dynamics misalignment between source and target domains.
Analyzes KL regularization estimators in RL training of LLMs, comparing bias-variance tradeoffs of different approximation methods.
VL-RouterBench benchmark for evaluating vision-language model routing systems with quality-cost tradeoff assessment at scale.
Evaluates feature-dependent noise in preference-based reinforcement learning with realistic noise patterns correlated to observations.
Proposes GIFT method reconciling SFT and RL post-training for Large Reasoning Models via Gibbs initialization to prevent distributional collapse.
Solves constrained optimization problems via gradient-based methods using hierarchical score-matching spaces to overcome local optima.
Proposes neural characteristic function approach for graph domain adaptation addressing distributional shifts without manual feature design.
Studies sequential prediction with option to abstain in semi-adversarial settings mixing adversarial and stochastic instances.
Creates deep surrogate model for blast wave prediction that generalizes to out-of-distribution urban scenarios using machine learning.
Develops federated causal representation learning for decentralized counterfactual reasoning across coupled industrial systems while preserving data privacy.
Proposes Hyperparameter Trajectory Inference to adjust neural network hyperparameters post-deployment without full retraining using optimal transport.
Studies how pretrained Vision-Language-Action models resist catastrophic forgetting during continual learning in robot policy training.
Reduces transformer KV cache by using low-dimensional keys for attention selection while maintaining high-dimensional values, achieving O(log N) dimensional compression.
JAWS improves neural PDE solvers' long-term rollouts using spatially-adaptive Jacobian regularization to prevent spectral blow-up and unphysical divergence.