RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
Adversarial attack on Mixture-of-Experts LLMs exploiting routing mechanisms to bypass safety alignment.
Adversarial attack on Mixture-of-Experts LLMs exploiting routing mechanisms to bypass safety alignment.
Kernel Affine Hull Machines enabling efficient query-side semantic encoding without repeated neural inference for transformer retrieval systems.
ZeRO-Prefill optimization reducing distributed execution overhead in mixture-of-experts model serving for prefill-only discriminative tasks.
Bridge diffusion method with closed-form solutions for score functions and drift fields enabling analytical controlled path generation without neural networks.
ISAAC framework auditing causal reasoning in deep learning drug-target interaction models via intervention-based structural sensitivity probing.
Benchmark for evaluating LLM agents with tool use detecting reward hacking exploits through multi-step task evaluation with naturalistic shortcuts.
Diffusion-aided reward shaping approach for scheduling AIGC workloads across distributed data centers while minimizing energy costs.
AutoRAGTuner declarative framework automating RAG pipeline optimization through modular architecture and configuration-driven hyperparameter tuning.
Gradient transport analysis framework for understanding cascade efficiency and transport mechanisms during large language model pretraining.
Cross-lingual safeguard transfer framework improving multilingual safety alignment in LLMs through self-distillation from high-resource languages.
Diffusion bridge framework for unpaired modality translation with structured constraints to restrict solution space.
TCD-Arena benchmark for evaluating robustness of time series causal discovery algorithms against violation of assumptions.
MechaRule framework extracting symbolic rules from LLMs by grounding rule extraction in neuron-level mechanistic interpretability via contrastive ablation.
Sample-efficient fine-tuning algorithm for diffusion and flow-based robot control policies using off-policy critic networks and modified PPO.
Adaptive negative sampling scheduler for graph contrastive learning improving computational efficiency in self-supervised representation learning.
Attribution-guided masking method for improving transformer sentiment classification transfer across domains by identifying spurious tokens.
Prompt arithmetic technique for improving LLM robustness under distribution shift without full fine-tuning via task arithmetic and prompt manipulation.
arXiv paper introducing Ensemble Directional Kalman Filter for pose tracking using unit-quaternion attitude representation beyond standard Kalman assumptions.
arXiv paper on Gated Subspace Inference accelerating transformer inference by exploiting low-rank activation manifolds with cached weights and per-token gating.
arXiv paper on distributionally robust Markov games for multi-agent RL addressing robustness and data efficiency in uncertain environments with large state spaces.
arXiv paper proposing Normalized Excess Cost (NEC) metric for classification evaluation that weights errors by per-example costs for safety-critical applications.
Benchmark for measuring classifier recovery speed under distribution shift and online corrections, addressing real-world deployment scenarios.
Pairwise matrix protocol for sparse autoencoder interpretability reveals limitations of standard single-feature inspection on language model feature analysis.
Framework for evaluating bias in LLMs through behavioral profiling and mechanistic interpretability, introducing Moral Sensitivity Index metric.
Theoretical analysis of activation alignment methods (RSA, CCA, CKA) for comparing neural representations in biological and artificial systems.
Self-Mined Hardness method for LLM safety fine-tuning that scores prompt difficulty by model jailbreak frequency, then trains on hardest prompts.
Text-Conditional JEPA uses image captions to reduce prediction uncertainty in masked visual representation learning.
Ortho-Hydra addresses style bleed in LoRA fine-tuning of diffusion transformers using orthogonalized mixture-of-experts architecture.
Examines whether LLMs exhibit core beliefsāfoundational commitments that resist changeāparallel to human cognition.
Investigates why transformers fail at counting tasks, finding failures stem from token generation rather than internal representation, with proposed fixes.
RFPrompt enables efficient adaptation of wireless foundation models to out-of-distribution modulation classification tasks via prompt-based expert adaptation.
Graph unlearning technique using feature-dimension aware quantile selection for privacy-preserving multimodal graph learning.
Studies convex optimization under adversarial gradient perturbations in distributed learning, analyzing privacy-utility tradeoffs.
DGPO algorithm for fine-grained credit assignment in LLM reinforcement learning, improving reasoning step isolation in chain-of-thought generations.
LLM-ADAM agent framework for detecting anomalies in additive manufacturing pre-prints, helping users without manufacturing expertise catch process-planning errors.
S3 framework decomposes multimodal inputs into semantic experts with selective routing and sparsification via mixture-of-experts approach.
Addresses imitation learning in mean-field games with stochastic population distribution, developing algorithms for population-aware agent policies.
Investigates zeroth-order optimization methods for fine-tuning large language models, explaining why they work despite theoretical dimension-dependent slowdown predictions.
Proposes GRAFT, a global explanation framework for Graph Neural Networks that identifies which input node attributes drive model predictions.
Studies how repeated LLM inference sampling improves accuracy beyond single-call performance, analyzing correctness correlation across examples through latent probability distributions.
Research shows differential privacy-perturbed GNN explanations can leak graph structure through adversarial reconstruction attacks.
Uses LLMs to automatically synthesize complete RL task interfaces including observations and rewards from natural language.
Proposes learning-to-theorize approach where agents build internal theories from observations inspired by cognitive science.
Introduces FIBER optimizer correcting bias from filtering in differentially private training with adaptive optimizers.
Proposes DynaTab, architecture with dynamic feature ordering for high-dimensional tabular data using neural rewiring.
Studies hierarchical reinforcement learning using variational quantum circuits.
Proposes AEMG, first large-scale self-supervised framework for learning generalizable EMG representations across subjects and devices.
Develops fast approximate online graph-based algorithms for semi-supervised learning and anomaly detection.
Proposes resolution-invariant diffusion models for function spaces over irregular domains using score-based approach.
Applies physics-informed neural networks with meta-learning to solve inverse problems in high-dimensional ODEs.