STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
Defense mechanism against adversarial patches in Vision Transformers using token segregation and randomized transformations.
Defense mechanism against adversarial patches in Vision Transformers using token segregation and randomized transformations.
Hierarchical LLM-based approach for fine-grained multi-table retrieval using compositional reasoning instead of coarse-grained similarity matching.
In-context learning strategy for CAD code generation using design-specification tiling to improve LLM performance on domain-specific tasks.
Foundation model-guided approach for virtual immunohistochemistry staining from H&E images to accelerate pathology diagnostics.
Multimodal recommendation framework using anchor-based alignment in projection space to prevent modality collapse and ID dominance.
Novel 3D molecule generation framework using vector-field representations to address modality entanglement and geometry-chemistry constraints.
Self-supervised system for robots to detect and recognize novel objects from human video demonstrations without prompt engineering.
Efficient optimization technique addressing long-tail distribution problem in LLM-based sequential recommender systems.
Multimodal foundation model for Earth observation using temporal training objectives robust to variable-length satellite and sensor data.
Method for injecting auxiliary visual features into vision-language-action models to improve geometric understanding and temporal reasoning for robotic manipulation.
Model distillation approach compressing 2B vision-language retriever into 70M text-only encoder for efficient document retrieval.
AI-based framework transforming global weather forecasts into fine-grained wind field predictions and infrastructure failure probabilities for tropical cyclones.
Research on using vision-language models for detecting and localizing forged images, studying how VLM priors affect forgery detection performance.
Data-efficient MRI reconstruction strategy using diffusion probabilistic models with pre-training and fine-tuning.
Adaptable fraud detection system handling adversarial attacks in resource-constrained environments with multiple risk modules.
Framework (ARL-Tangram) optimizing resource efficiency in agentic RL by dynamically allocating external compute resources.
Surgical world model using controllable video generation for simulating surgical actions with precise tool-tissue control.
Dataset and baseline for real-time screw classification in industrial automation and robotic systems.
Diagnostic benchmark (ESPIRE) for evaluating vision-language models on embodied spatial reasoning tasks.
GNN approach for precoder learning in cell-free wireless systems accounting for dynamic user-access point associations.
Empirical study of federated few-shot learning on neuromorphic hardware using spike-timing-dependent plasticity.
Convergence analysis of functional learning methods for contextual stochastic optimization problems.
Research on interpretable multimodal concept bottleneck models ensuring faithful explanations through proper concept detection.
Provable multi-agent reinforcement learning in partially observable stochastic games leveraging information sharing among agents.
Introduces diffusion models as expressive variational posteriors for black-box inference in latent variable models.
Graph signal processing research extending sampling theory to graphon signals using limits of large graphs.
Research on offline reinforcement learning combining return-conditioned supervised learning with Q-functions to improve stitching capability and stability.
Introduces Walk Profile method and explores positional encodings for directed graphs in graph neural networks and graph transformers.
Explores polynomial attention alternatives to softmax in transformers, arguing regularization rather than probability distribution drives performance.
Proposes first computationally efficient algorithm with optimal regret for infinite-horizon discounted reinforcement and imitation learning.
Advocates integrating causal methods into ML to balance trustworthiness objectives like fairness, privacy, robustness, and explainability.
Proposes Dual Filter framework connecting Hidden Markov Models to transformer decoder architecture for causal nonlinear prediction.
Introduces Guided Policy Optimization framework for RL in partially observable environments using privileged information from simulators.
Analyzes oversmoothing problem in deep Graph Neural Networks and explores why networks fail to learn non-oversmoothed representations.
Proposes uncertainty estimation improvements to Residual Reinforcement Learning for faster adaptation of pretrained policies with sparse rewards.
Data condensation approach for training diffusion models with minimal computational budget by constructing smaller synthetic training datasets.
Graph transformer architecture designed for invariant learning to improve out-of-distribution generalization on graph-structured data.
Theoretical study of implicit bias in deep neural network training showing gradient flow induces learning of lower-dimensional parameter structures.
Continual learning framework with unified prompt pools for medical imaging tasks, addressing domain-specific challenges in adaptive AI.
Analysis of compositional generalization mechanisms in conditional diffusion models, studying length generalization on controlled image generation tasks.
Lightweight meta-learning method using three parameters to dynamically adjust sample loss weights for noisy training, fairness, and synthetic data utilization.
Hybrid pre-training approach using low-rank adapters alongside full training to reduce computational cost for vision transformer training.
Method for robust fine-tuning non-robust pretrained models using epsilon-scheduling to achieve adversarial robustness and task adaptation simultaneously.
Analysis of transformer internals distinguishing recall from reasoning mechanisms through layer-wise attention and activation patterns for interpretability.
Mathematical proof that transformer language models are injective, enabling exact input recovery from representations despite nonlinear components.
Method for unlearning harmful content from LLMs by analyzing belief redistribution in probability space, avoiding unwanted side effects of gradient ascent.
Theoretical analysis of data scaling laws in linear regression when training multiple epochs on limited datasets, relevant to LLM training efficiency.
Method to improve LLM consistency and reliability across semantically equivalent prompts using group relative policy optimization for business-critical applications.
Study demonstrating that ensemble diversity across language models mitigates knowledge collapse from training on model-generated outputs.
Mixture-of-experts approach with heterogeneous experts for capturing multi-scale temporal dynamics in long-horizon time series forecasting.