Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection
KV cache optimization using low-dimensional attention selection to reduce transformer memory with O(log N) key dimensions.
KV cache optimization using low-dimensional attention selection to reduce transformer memory with O(log N) key dimensions.
Study examining many-shot prompting for test-time LLM adaptation, analyzing reliability and limits of in-context learning scaling.
Active feature acquisition method for biomedical applications optimizing measurement selection under temporal and cost constraints.
Theoretical analysis proving attention sinks are functionally necessary in softmax transformers for certain tasks.
Research addressing multimodal model underperformance in context-aided forecasting via improved context quality assessment.
Machine learning method using hypergraph pre-training to improve atrial fibrillation prediction in stroke patients.
Interactive benchmark environment for synthesizing flat-foldable origamis, testing AI systems' planning and causal reasoning in physical domains.
Neural compression framework using SIREN auto-decoders for high-fidelity compression of multi-structural seismic velocity models.
Graph-based verifier for LLM task planning that identifies and corrects hallucinations and flaws in agent-generated plans.
Post-hoc model-agnostic explanation method using informative perturbation selection for uncertainty-aware interpretability of black-box ML models.
Convergence analysis of Muon optimizer under heavy-tailed noise for nonconvex optimization in large-scale deep neural network training.
Analysis showing wider beam search in LLMs can degrade output quality due to overestimation bias in noisy scorer outputs, with theoretical grounding.
Large-scale benchmark for AI agents combining partial observability, game-theoretic reasoning, and long-horizon planning in Pokemon battle environment.
Physics-informed neural networks and neural operators for simulating EUV electromagnetic wave diffraction in lithography mask applications.
Annotation-free method for reconstructing controllable 3D Gaussian splats of articulated objects from monocular video using flow derivatives.
Data synthesis engine using scene graphs to improve compositional generalization and semantic alignment in text-to-image generation models.
CHARM method calibrating reward models using Chatbot Arena scores to mitigate model preference bias, improving alignment of LLMs through RLHF.
PhysioOmni foundation model for multimodal physiological signals handling arbitrary missing modalities across EEG, ECG, EOG, EMG for healthcare and brain-computer interfaces.
BiomedSQL benchmark for text-to-SQL generation requiring scientific reasoning over biomedical knowledge bases, evaluating LLM capability for complex analytical tasks.
DP-Powered LLMs for privacy-preserving radiology report classification, enabling differential privacy in healthcare diagnosis and abnormality classification workflows.
TempCore benchmark analyzing whether video QA models genuinely require temporal frame selection, introducing Frame Selection Sensitivity metric for VLM diagnostic evaluation.
Comparison of statistical and logic-based XAI techniques for interpreting ML security alerts in 5G intrusion detection systems, enabling actionable incident response.
ERGO framework for efficient high-resolution image processing in vision-language models using coarse-to-fine reasoning pipeline with two-stage visual token reduction.
Hilbert system recursively building formal proofs by combining informal LLM reasoning with Lean 4 verification, bridging gap between mathematical reasoning and formal proof generation.
Zephyrus framework combining foundation models for weather forecasting with LLM reasoning to enable language-based scientific workflows on meteorological datasets.
Security research on backdoor attacks in AI agent supply chains through poisoned interaction data collection, formalizing threat models for finetuned web browsing and tool-use agents.
Training-free diffusion model enabling layer-wise control in text-to-image generation through noise transplantation without fine-tuning or large datasets.
Deep Q-Network learns satellite weighting for CSI-free multi-satellite positioning in LEO constellations combined with weighted least squares estimation.
Topological data analysis patch-based approach for CT imaging feature extraction improving ML model performance on medical diagnosis tasks.
Noise-aware masked autoencoder for self-supervised SAR satellite imagery representation learning addressing data scarcity and speckle noise challenges.
FusionRoute enables token-level collaboration between specialized and general-purpose LLMs via dynamic routing, improving efficiency and domain performance.
Language-aligned concept foundation model decomposing vision representations into human-interpretable concepts with spatial grounding across diverse tasks.
VisTIRA addresses vision-language model performance gap on visual math problems through structured tool integration and dense formula/layout handling.
Linear probes on LLM pre-generation activations predict task success before inference, enabling efficient routing and reducing extended reasoning compute costs.
Addresses cross-agent noise in multi-agent reinforcement learning through descent-guided policy gradients, reducing sample complexity from O(N/ε) to improved bounds.
Proposes Model Medicine framework for understanding, diagnosing, and treating disorders in AI models using clinical methodology analogies for interpretability.
Evaluates vision foundation models adapted for pasture biomass regression, comparing SSMs, transformers, and simpler cross-view modules on agricultural imagery.
Survey introducing reinforcement learning methods to economists for solving high-dimensional dynamic programming problems that resist dimensionality reduction.
Framework defining AI models and AI systems through systematic review of 896 academic papers and 80+ regulatory documents to resolve boundary problem in AI regulation.
Red-teaming study on adversarial activation steering techniques to compromise LLM responses, examining safety risks from semantic layer manipulation.
Evaluates autonomous cyber-attack capabilities of frontier AI models on multi-step attack scenarios, comparing seven models over 18 months at varying inference compute budgets.
Gradient flow utility metric for structural pruning and dynamic routing in deep networks, addressing magnitude bias in weight-based heuristics.
Foundation models as surrogates for active learning in materials discovery, reducing costly synthesis cycles through better uncertainty estimation.
Empirical comparison of CNNs, contrastive VLMs, and generative VLMs for crop disease classification across diverse conditions.
SHAMISA framework for no-reference image quality assessment using self-supervised learning on unlabeled distorted images.
Data augmentation method for ring-type polygon annotations in floorplan analysis preserving topology during geometric transformations.
Self-supervised learning framework using masked BRep autoencoder with hierarchical graph transformer for CAD model representation learning.
Research analyzing error sources in global feature effects (PD/ALE plots) for black-box model interpretation.
HindSight framework evaluates LLM-generated research ideas using time-split evaluation against future publications and citation impact.
GSD 2 is a standalone CLI coding agent built on Pi SDK, evolving from a Claude prompt framework to a full agent with session and context control.