CERSA: Cumulative Energy-Retaining Subspace Adaptation for Memory-Efficient Fine-Tuning
CERSA parameter-efficient fine-tuning method reducing memory constraints by capturing full rank characteristics of weight updates.
CERSA parameter-efficient fine-tuning method reducing memory constraints by capturing full rank characteristics of weight updates.
Echo-LoRA parameter-efficient fine-tuning technique injecting cross-layer representations to improve LLM adaptation with lower rank updates.
Federated graph learning framework for discovering novel categories in decentralized graph data without closed-world assumptions.
Introduces Information Density metric to optimize sensor deployment and enable AI-driven virtual sensing in IoT networks.
Parameter-efficient fine-tuning method operating in frequency domain with dynamic text-guided adaptation for pre-trained models.
Proposes RQIQN, a Wasserstein distributionally robust enhancement for quantile-based distributional reinforcement learning.
Analyzes attention mechanisms in multimodal transformers through neuroscience lens to understand visual interestingness encoding.
Characterizes normalization equivariance in image-to-image prediction networks for improved robustness to distribution shift.
Benchmark of 1,300 items for evaluating causal mechanism induction from interventions using executable Boolean DSL.
Open-source Python library for fairness auditing, privacy preservation, and trustworthy ML in healthcare applications, especially low-resource settings.
Neurosymbolic framework combining neural networks with symbolic reasoning for visual understanding using weak supervision.
Proposes diffusion-based method for detecting out-of-distribution actions in offline reinforcement learning without penalization-based suppression.
Fine-tuning Whisper and PyAnnote models for Bangla speech recognition and speaker diarization on long-form recordings.
Inference-time method improving LLM reliability by injecting controlled noise into latent representations to generate diverse counterfactual outputs.
Joint optimization of approximate multiplier structures and AI model training for power-efficient accelerators.
Diagnostic framework for KV cache compression in long-context LLM inference identifying why value-aware eviction selectors succeed or fail.
Self-supervised learning approach for microcontroller models under 500K parameters using teacher-guided distillation.
Mechanistic analysis of hallucination in vision-language models caused by geometric over-alignment between visual and language representations.
LLM-based translation of compiler intermediate representations between GCC and LLVM for cross-toolchain interoperability.
Study of LLM performance on predicting polymer physical properties from synthesis and processing descriptions.
Security framework for adversarial robustness in medical decision-making LLM agents with full-link enhancement pipeline.
Methodological analysis exposing benchmarking flaws in computer use agents, showing static evaluation metrics fail to measure genuine interactive reasoning.
Standardized admission contract framework for heterogeneous AI execution requests including inference, evaluation, and agentic workflows in enterprise systems.
Analysis of security vulnerabilities in multi-agent LLM systems where malicious agents exploit consensus formation mechanisms.
Memory-augmented agentic approach for understanding ultra-long videos using multimodal LLMs with structured memory retrieval across modalities.
Graph neural network approach for traffic forecasting using prompt learning to improve generalization across spatio-temporal distribution shifts.
Technique to mitigate many-shot jailbreak attacks on safety-aligned LLMs by detecting progressive activation drift from harmful demonstrations.
Defense mechanism against backdoor attacks on Graph Neural Networks by analyzing trigger correlations and dependencies rather than surface-level patterns.
World models for embodied AI that learn predictive models from visual observations with physical consistency constraints for reinforcement learning and robotic planning.
HTPO: Hierarchical token-level RL objective control for balanced exploration-exploitation in LLM reasoning via Chain-of-Thought.
Task-aligned GNN analysis for Electronic Design Automation, aligning graph propagation with native task algebra.
Hi-MoE: Hierarchical Mixture-of-Experts with two-stage routing optimization balancing load distribution and expert specialization.
In-Context Fixation: LLM label-slot fixation phenomenon where homogeneous demonstration labels collapse few-shot classification accuracy across models.
Theoretical analysis of scaling behavior in normalized residual networks through depth expansion and test-risk mechanisms.
Code retrieval via embedding-based methods using LLM rephrasing strategies at varying levels: stylistic, pseudo-code, and natural language transcription.
mHC-SSM: Manifold-constrained hyper-connections for State Space language models with stream-specialized adapters.
Priming: Method to initialize hybrid State-Space models from pre-trained Transformers, achieving faster decoding and smaller caches.
LLMSYS-HPOBench: Hyperparameter optimization benchmark for LLM systems addressing compound configuration spaces and AutoML challenges.
WebTrap: Security vulnerabilities in browser agents exploited through mid-task prompt injection attacks during navigation, exposing gaps in real-world deployment robustness.
SeedHijack demonstrates a backdoor attack exploiting PRNG manipulation in LLM sampling and proposes quantum RNG defense mechanisms.
FlashSVD v1.5 is a runtime system for efficiently serving SVD-compressed transformers by addressing fragmentation overhead in prefill and decode phases.
RDKV uses rate-distortion bit allocation to optimize LLM inference with long contexts by managing KV cache eviction and quantization.
Multi-Scale Attention Transformer architecture for solving PDEs on irregular domains; compares attention vs Fourier-based operators.
Mazocarta is an instrumented procedural deckbuilder in Rust/WebAssembly supporting both interactive play and automated simulation.
LLM Wardens: secondary LLM monitors conversations to protect users from adversarial persuasion; user study shows 65% success rate without intervention.
SDG-MoE architecture enables communication among routed experts in sparse Mixture-of-Experts models for improved performance.
Federated quantum neural network for privacy-preserving diabetic retinopathy detection from medical images.
CAMAL method uses segmentation masks to improve attention alignment and faithfulness in vision transformers.
Research paper on monetizing LLMs through generative advertising using neuron auction mechanisms balancing revenue and user experience.
DPA-GRPO training method for reliable structured LLM generation ensuring local correctness, global consistency, and auditability in decision workflows.