Causal Bandit Over Unknown Graphs: Upper Confidence Bounds With Backdoor Adjustment
Paper on causal bandit algorithms for unknown DAGs using confidence bounds and backdoor adjustment for intervention discovery.
Paper on causal bandit algorithms for unknown DAGs using confidence bounds and backdoor adjustment for intervention discovery.
Research on finite-horizon restless bandit problems reformulated as thresholding with improved sample complexity and policy convergence.
arXiv paper on decentralized learning using consensus gradient descent with privacy and communication constraints across networked devices.
Research paper on model stealing attacks and defenses, analyzing vulnerabilities of ML services to adversarial extraction through query access.
Analytical framework explaining spectral bias in diffusion model training dynamics using Gaussian equivalence and probability-flow ODEs.
RaPA improves transferable targeted adversarial attacks by random parameter pruning to reduce reliance on surrogate model subsets.
Finite-time convergence analysis for average-reward Q-learning with adaptive stepsizes, showing O(1/k) convergence rate.
First mechanistic interpretability framework for VAEs using multi-level causal interventions to understand generative model representations.
Introduces Bayesian ablation framework for interpreting learned task representations in neural networks through probabilistic inference.
MSDformer extends discrete token modeling for time series generation using multi-scale transformer architecture to capture temporal patterns.
SoSBench benchmarks safety alignment of LLMs across six scientific domains with sophisticated risks beyond basic misuse scenarios.
Studies the problem of using LLMs as judges for evaluating LLM outputs, addressing epistemic uncertainty in judge quality beyond sampling variability.
K-Steering enables unified multi-attribute control of LLMs at inference time using non-linear classifiers on hidden activations to handle attribute interference.
MLorc proposes momentum low-rank compression for memory-efficient LLM fine-tuning, reducing memory demands compared to LoRA while maintaining performance.
SFBD Flow framework trains diffusion models on corrupted/noisy data with clean samples to reduce privacy risks and improve convergence in generative modeling.
Token significance approach in RL for efficient LLM reasoning by identifying and prioritizing important tokens over length optimization.
Federated Item Response Theory (FedIRT) framework enabling distributed psychometric estimation without centralizing raw response data.
Causal Process Models for learning sparse time-varying causal graphs from visual observations using reinforcement learning.
Causal multi-armed bandit algorithm reasoning under uncertain causal mechanisms from graphical models.
xRFM feature learning models for tabular data providing accurate, scalable, and interpretable alternatives to gradient boosted trees.
LoFT parameter-efficient fine-tuning approach for long-tailed semi-supervised learning leveraging foundation models.
Agentic Classification Tree (ACT) combining LLMs with decision trees for transparent, interpretable decisions on unstructured data.
Security analysis exposing vulnerabilities in LLM weight pruning methods used by inference engines like vLLM.
Theoretical finite-time analysis of Q-learning with time-varying policies under minimal assumptions for Markov decision processes.
Self-evolving Post-Training (SePT) method enabling LLMs to improve reasoning without external rewards through self-generated training data.
Information-theoretic analysis of out-of-distribution generalization in meta-reinforcement learning with bounds under distribution shift scenarios.
Classical clustering algorithm handling arbitrary cluster geometry without global density assumptions using skeleton propagation and recalibrating expansions.
Method for ranking synthetic datasets by real-world performance without annotations, establishing benchmarks for synthetic data quality estimation.
Feynman-Kac framework for guiding diffusion-based generative models toward proteins with specified properties and tailored structures.
Unified stability analysis comparing SAM and SGD optimization algorithms showing role of data coherence and simplicity bias in generalization.
Comparative analysis of transformer models (DistilBERT) versus psycholinguistic features for detecting business email compromise attacks.
Framework addressing structural overfitting in graph neural networks for missing feature imputation using distribution-aware rectification.
Federated learning approach for vehicle edge caching using personalized distillation to predict user content preferences while preserving privacy.
Random-bridges framework for generative models using stochastic processes conditioned on target distributions for flexible transport between distributions.
Electric load forecasting model integrating multi-source textual data (news, social media, policies) with temporal grid-aware predictions.
Mamba-based neural operator framework for accurate chemical kinetics modeling in combustion simulations using efficient temporal modeling.
Distribution restoration method using noisy samples and optimal transport to recover fully observed data from partial corrupted observations.
Method using large language models to measure semantic similarity in categorical data clustering by bridging gap in attribute distance representation.
Differentiable adversarial framework for task-aware data reduction using learnable selector and minimax optimization to identify informative samples.
LLM-based hardware-aware quantization agent automating model quantization for efficient LLM deployment on resource-constrained hardware.
Theoretical analysis of implicit bias in stochastic learning using geometric perspective to explain solution selection in overparameterized models.
Split learning optimization method reducing memory overhead for LLM training on edge devices using hybrid-order optimization instead of first-order approaches.
Neural memory storage architecture for LLMs with invertible compression and learnable prediction for runtime memory.
Information-theoretic approach for designing shared visual tokenizers in unified multimodal LLMs.
Safety alignment framework addressing unique challenges of sparse routing in Mixture-of-Experts language models.
Training framework for multimodal systems to maintain performance when input channels are lost at deployment.
Pruning framework for physics-informed neural networks to improve robustness to noise in PDE inverse problems.
Time series imputation method using channel-head binding for handling diverse missing patterns.
Training method for LLMs to directly model generative reasoning process in scientific discovery applications.
Multi-agent reinforcement learning algorithm for general-sum games with convergence guarantees in heterogeneous agent settings.